Updated
Updated · Barchart · Aug 15
AI Firms Destroy Millions of Books for Training Data as $1.5 Billion Copyright Costs Loom
Updated
Updated · Barchart · Aug 15

AI Firms Destroy Millions of Books for Training Data as $1.5 Billion Copyright Costs Loom

2 articles · Updated · Barchart · Aug 15

Summary

  • Booksellers in Australia, Europe and the UK are seeing bulk purchases of obscure and out-of-print titles that are being scanned for AI training and then pulped rather than resold or archived.
  • Anthropic court filings show the practice at scale: under "Project Panama," it bought copyrighted books, cut off bindings for easier scanning and destroyed many physical copies after digitizing them.
  • The buying spree reflects a scramble for higher-quality, pre-2022 human-written text as freely available web data becomes less useful and more contaminated by AI-generated material.
  • A $1.5 billion Anthropic settlement with authors and a similar lawsuit against Google highlight how training-data costs could rise sharply, pressuring AI economics even if some courts keep treating the copying as fair use.
  • That leaves a dual risk beyond copyright fights: rare nonfiction and academic works can disappear from circulation, while the price of scarce, reliable training data keeps climbing for AI developers and investors.

Insights

Could the sudden disappearance of random secondhand books signal a desperate, covert data war among global tech companies?
Why are mystery buyers secretly hoarding and destroying thousands of obscure physical books worldwide?