AI Firms Destroy Millions of Books for Training Data as $1.5 Billion Copyright Costs Loom
Updated
Updated · Barchart · Aug 15
AI Firms Destroy Millions of Books for Training Data as $1.5 Billion Copyright Costs Loom
2 articles · Updated · Barchart · Aug 15
Summary
Booksellers in Australia, Europe and the UK are seeing bulk purchases of obscure and out-of-print titles that are being scanned for AI training and then pulped rather than resold or archived.
Anthropic court filings show the practice at scale: under "Project Panama," it bought copyrighted books, cut off bindings for easier scanning and destroyed many physical copies after digitizing them.
The buying spree reflects a scramble for higher-quality, pre-2022 human-written text as freely available web data becomes less useful and more contaminated by AI-generated material.
A $1.5 billion Anthropic settlement with authors and a similar lawsuit against Google highlight how training-data costs could rise sharply, pressuring AI economics even if some courts keep treating the copying as fair use.
That leaves a dual risk beyond copyright fights: rare nonfiction and academic works can disappear from circulation, while the price of scarce, reliable training data keeps climbing for AI developers and investors.