Training data sourcing and licensing, data quality, and synthetic data.

Current topic Data & Training · 20 Switch topic

Latest stories

  1. Google figures out how to watermark AI-designed proteins

    Google DeepMind published a paper on SynthID Bio, which adds watermarks to AI-designed proteins without harming their function. AI-designed proteins are hard to catch with existing DNA synthesis screening for dangerous sequences. DeepMind modified the widely used design tool ProteinMPNN to embed a signal in amino acid choices, and wet-lab tests on VEGF-A, the SARS-CoV-2 spike RBD and PD-L1 showed hit rates and binding affinity matching unwatermarked designs. It also built the capability into AlphaFold 3.

    View in daily briefing Read original

  2. Meta's Muse Is Better at Surveilling Than Helping Me

    Meta's new AI agent app Muse was downloaded over 900,000 times in its first week, but turned out to be more eager to keep collecting personal data, like bank accounts and email, than to actually get things done. Muse can browse the web on its own to run errands like ordering from a bakery or buying used furniture, but it also pushes users to add a payment card and permanently stores chat details in a 'Memory' file with no way to turn that feature off entirely.

    View in daily briefing Read original

  3. China Suggests Anthropic Could Hand User Data to US Intelligence

    A CCTV-affiliated social media account, Yuyuan Tantian, took aim at Anthropic's revised privacy policy on September 19. It claimed Anthropic has amended its policy 13 times since 2023, changing terms so that data from users in Canada, Brazil, South Korea, and the EU could be shared with US intelligence agencies without legal process. The claim, however, comes from a Chinese state-media-linked account ahead of a US-China summit and remains independently unverified.

    View in daily briefing Read original

  4. UN turns to Google to make its global data ready for AI agents

    The United Nations launched the UN System Data Commons with Google, restructuring its statistics so AI agents can query and use them directly. Built on Google's open-source Data Commons platform, it answers natural-language questions and lets AI agents connect through the Model Context Protocol (MCP) standard. The move follows a UNICEF benchmark that tested six large language models from OpenAI, Anthropic, and Google on more than 133,000 global development questions and found an average accuracy of just 21.2 percent, with about 60 percent of answers failing to return a usable number at all, and only about half of repeated questions returning the same figure two days later. AI-driven traffic to UNICEF's data site is already substantial: visits arriving via links in ChatGPT answers rose 67 percent year over year and now account for 6.4 percent of all sessions. Twenty-six UN agencies have committed to the platform, about 20 have data live at launch, and the UN aims to move 80 percent of its statistics onto it by 2027, with Google.org providing $2 million in funding and technical support.

    View in daily briefing Read original

  5. Virtual fruit fly brain simulation sparks unusual experiments, even a 'fly language model'

    A joint team from Google Research and the Howard Hughes Medical Institute's Janelia campus published the first complete simulated neural network of an adult male fruit fly on September 3. The simulation includes 166,700 virtual neurons and roughly 25 million synaptic connections, and reacts to external stimuli the way a real fly would. Developer Alex Wormus connected the network to a 1.2-billion-parameter language model to build a "Fly Language Model" (FLM), but found that the biological network didn't meaningfully improve the model's language abilities.

    View in daily briefing Read original

  6. Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data

    Mecka AI, a startup that pays people to record everyday motions like making coffee or fixing a car for humanoid robot training, is nearing a new funding round led by Sequoia Capital at a roughly $500 million valuation, TechCrunch reports. That's just three months after it closed a $60 million Series A led by Framework Ventures. The company, founded in 2024 by four people with no robotics background, is aiming for $100 million in annual run rate (ARR) by the end of 2026.

    View in daily briefing Read original

  7. OpenAI Brings a 'Data Agent' for Enterprise Analytics to ChatGPT Work

    OpenAI added a new 'Data Agent' to ChatGPT Work on September 10, letting employees analyze internal company data using plain-language questions like 'why did users drop last week' without learning a specialized analytics tool or writing queries. It connects to databases such as Amazon Redshift, Databricks, Snowflake, and MongoDB alongside documents from Google Drive and SharePoint, and it integrates with existing BI tools like Tableau and Power BI.

    View in daily briefing Read original

  8. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

    Google DeepMind released AlphaGenome Atlas on September 8, predicting the molecular impact of roughly 9 billion possible single-nucleotide variants across the human genome. The dataset spans 1 petabyte, more than 30 times larger than the AlphaFold Database. Its 'AVI score' reduces each variant's impact to a single number, covering not just the 2 percent of the genome that codes for genes but the remaining 98 percent of non-coding regions too.

    View in daily briefing Read original

  9. Data from drones in Ukraine is fueling a new Wild West marketplace

    Ukraine's defense ministry has, since January, opened footage from hundreds of thousands of drone flights to more than 100 defense and commercial companies as well as the UK government, creating a new market for battlefield data. US firm Enabled Intelligence has already processed over 500,000 hours of Ukrainian drone footage for training commercial and military AI models. The data is especially valuable because it captures messy, lab-unreplicable conditions like signal jamming and lost visibility.

    View in daily briefing Read original

  10. Meta is paying to peek at how you use their latest AI model

    Meta is pricing its new coding-agent model, Muse Spark, differently based on whether developers agree to share their data. Developers who let Meta use their prompts and outputs for training pay $0.10 per million input tokens instead of $1.25, and $0.20 per million output tokens instead of $4.25, cuts of more than 90%. Princeton professor Arvind Narayanan noted that large companies tend to skip this discount and pay more for enterprise plans instead, because of data retention and governance concerns.

    View in daily briefing Read original

September 20261

August 20269