Updated
Updated · Kotaku · Aug 1
ISBNdb Pulls AI Book-Training Page as $1.5 Billion Anthropic Case Fuels Rare-Book Backlash
Updated
Updated · Kotaku · Aug 1

ISBNdb Pulls AI Book-Training Page as $1.5 Billion Anthropic Case Fuels Rare-Book Backlash

3 articles · Updated · Kotaku · Aug 1

Summary

  • ISBNdb said it removed a marketing page that appeared to connect AI firms with printed books for training, insisting it has never trained models and never launched such a service.
  • The retreat followed outrage over reports that suppliers were getting unusual bulk orders for books that could be scanned and destroyed to create cleaner training data while avoiding copyright risks.
  • A deleted ISBNdb pitch had called books "the world's best AI training data," reinforcing suspicions after 404 Media reported on sellers unloading uncommon titles despite discomfort with the books being pulped.
  • The scrutiny widened after unsealed Anthropic "Project Panama" filings showed book scanning and destruction was treated as legally permissible; Anthropic later settled with authors for $1.5 billion.
  • The episode underscores a broader AI-data squeeze as companies hunt higher-quality, non-web sources and may favor destructive scanning simply because it is faster and cheaper.

Insights

Did a major metadata provider actually build a secret pipeline to feed physical books to AI, or was it truly just a market test?
Why are mysterious buyers bulk-ordering and destroying thousands of physical books, and what does this mean for the future of human-written literature?