Training data sourcing and licensing, data quality, and synthetic data.
-
Hidden Airtag reveals Amazon is trashing rare books to train AI
Investigative outlet 404 Media reported on August 17 that an AirTag hidden in a rare book by a cooperating bookseller led to Amazon's AI training facility, VGT3, in Las Vegas. Workers there tear books from their spines to scan the pages before discarding them, and the warehouse door even displays a logo of a T. rex devouring a book. Amazon's only response was a boilerplate statement that it "purchases books through commercial channels to help develop and improve the products and services our customers use."
-
Amazon, which started off selling books, is destroying rare texts to train AI
TechCrunch reported on August 17, citing 404 Media's investigation, that Amazon — once an online bookseller — is now destroying rare books to train its AI models. With freely available online text already exhausted, rare books offer AI firms unique training data that competitors can't easily access. Publications from before 2022 are also valued because they predate widespread AI-generated text online, helping models avoid model collapse.
-
As public data runs dry, companies' internal Slack, email, and recordings command a premium
With AI companies having largely exhausted publicly available internet data, they are now buying up internal corporate records — Slack messages, meeting recordings, emails, and code-change logs — as a new training-data source. Data broker Protege reports at least $100M in deals this year, more than triple last year's $30M, taking a 25-30% commission. AI agent startup Warmly received a $300,000 offer from data firm Mercor for its Slack, GitHub, Google Drive, and meeting transcripts just eight days after agreeing to be acquired by HubSpot, but turned it down.
-
Twitch content has trained Amazon AI for years, but users can opt out now
Twitch rolled out a new opt-out setting letting streamers and viewers refuse to have their content used to train Amazon's AI. Streams, VODs, clips, chat logs, and channel images are all eligible for training, and users remain opted in by default unless they turn it off. Amazon has used Twitch content for AI training for at least two years since acquiring the platform in 2014, a fact only confirmed publicly in 2024 by an executive's remark.
-
Scaling AI agents with trustworthy data
For AI agents to act autonomously, they need access to data across the whole company, but a survey of 300 data and technology executives run with Google Cloud found AI could reach only 45% of enterprise data on average. Among 'data leader' companies with over 70% access, 100% trusted their agents' decisions, while overall trust across companies sat around 50%. Sixty-six percent of 'data laggard' companies said legacy systems were blocking them from scaling agents.
-
The web’s newest weapon against AI scrapers is a font
A font called ShieldFont, built by two designers, uses data poisoning to defend against AI scrapers: it shows readers a normal page while feeding scrapers a version with word meanings swapped out. Using typographic ligatures, it replaces an average of 24.5% of all words and 45.8% of content words with unrelated terms, and in tests against six real scraper pipelines, over 90% of altered pages were rejected by quality filters. The defense breaks down, though, if a scraper simply renders the page as an image and reads it with optical character recognition.