Training data sourcing and licensing, data quality, and synthetic data.

Current topic Data & Training · 16 Switch topic

Latest stories

  1. Virtual fruit fly brain simulation sparks unusual experiments, even a 'fly language model'

    A joint team from Google Research and the Howard Hughes Medical Institute's Janelia campus published the first complete simulated neural network of an adult male fruit fly on September 3. The simulation includes 166,700 virtual neurons and roughly 25 million synaptic connections, and reacts to external stimuli the way a real fly would. Developer Alex Wormus connected the network to a 1.2-billion-parameter language model to build a "Fly Language Model" (FLM), but found that the biological network didn't meaningfully improve the model's language abilities.

    View in daily briefing Read original

  2. Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data

    Mecka AI, a startup that pays people to record everyday motions like making coffee or fixing a car for humanoid robot training, is nearing a new funding round led by Sequoia Capital at a roughly $500 million valuation, TechCrunch reports. That's just three months after it closed a $60 million Series A led by Framework Ventures. The company, founded in 2024 by four people with no robotics background, is aiming for $100 million in annual run rate (ARR) by the end of 2026.

    View in daily briefing Read original

  3. OpenAI Brings a 'Data Agent' for Enterprise Analytics to ChatGPT Work

    OpenAI added a new 'Data Agent' to ChatGPT Work on September 10, letting employees analyze internal company data using plain-language questions like 'why did users drop last week' without learning a specialized analytics tool or writing queries. It connects to databases such as Amazon Redshift, Databricks, Snowflake, and MongoDB alongside documents from Google Drive and SharePoint, and it integrates with existing BI tools like Tableau and Power BI.

    View in daily briefing Read original

  4. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

    Google DeepMind released AlphaGenome Atlas on September 8, predicting the molecular impact of roughly 9 billion possible single-nucleotide variants across the human genome. The dataset spans 1 petabyte, more than 30 times larger than the AlphaFold Database. Its 'AVI score' reduces each variant's impact to a single number, covering not just the 2 percent of the genome that codes for genes but the remaining 98 percent of non-coding regions too.

    View in daily briefing Read original

  5. Data from drones in Ukraine is fueling a new Wild West marketplace

    Ukraine's defense ministry has, since January, opened footage from hundreds of thousands of drone flights to more than 100 defense and commercial companies as well as the UK government, creating a new market for battlefield data. US firm Enabled Intelligence has already processed over 500,000 hours of Ukrainian drone footage for training commercial and military AI models. The data is especially valuable because it captures messy, lab-unreplicable conditions like signal jamming and lost visibility.

    View in daily briefing Read original

  6. Meta is paying to peek at how you use their latest AI model

    Meta is pricing its new coding-agent model, Muse Spark, differently based on whether developers agree to share their data. Developers who let Meta use their prompts and outputs for training pay $0.10 per million input tokens instead of $1.25, and $0.20 per million output tokens instead of $4.25, cuts of more than 90%. Princeton professor Arvind Narayanan noted that large companies tend to skip this discount and pay more for enterprise plans instead, because of data retention and governance concerns.

    View in daily briefing Read original

  7. Google needs Hollywood more than the studios need AI

    Google has been approaching major Hollywood studios, including Disney, Warner Bros. Discovery, and Universal, with licensing deals that would let it train AI models on their copyrighted libraries, according to the Los Angeles Times. Sources say Google could pay a studio like Disney around $40 million just for the rights to generate a single copyrighted character. No studio has agreed to a deal yet.

    View in daily briefing Read original

  8. Sony and Warner sue Anthropic over "one of the largest and most blatant ongoing thefts of intellectual property in history"

    Sony Music, Warner Music, and other music publishers have sued Anthropic in federal court in Northern California. The complaint alleges CEO Dario Amodei and co-founder Benjamin Mann personally directed and approved unauthorized downloads via torrenting, and that plaintiffs' songs, lyrics, and sheet music, tens of thousands of works in all, were taken from licensing platforms like Musixmatch and LyricFind as well as illegal digital archives.

    View in daily briefing Read original

  9. Spirit Airlines Wants to Sell Its Data to Google. Former Flight Attendants Are Freaked Out

    Google has agreed to pay $10 million for 34 years of bankrupt Spirit Airlines' operating data, including invoices, flight records, Wi-Fi sales, employee records, and crew pairings, beating a competing $7.5 million bid from AI data firm Mercor. Google says the deal excludes customer data and that it won't receive any personal information, but a union representing 5,500 former Spirit flight attendants has filed a court objection over the sale of employee data for AI training.

    View in daily briefing Read original

  10. Flight attendants freaked out that Google is buying tons of Spirit employee data

    Google won an auction to buy Spirit Airlines' employee dataset for $10 million, saying it will help train its AI models. That's $2.5 million more than the $7.5 million second-place bid from Mercor. The dataset includes roughly 100 million employee emails and a decade of HR, payroll, and activity records, and Google agreed to have a court-appointed third party strip personal identifiers before it takes possession.

    View in daily briefing Read original

August 20266