Updated
Updated · Ars Technica · Aug 12
AI Companies Destroy Books for Model Training as Google’s 2009 Scan Method Sits Unused
Updated
Updated · Ars Technica · Aug 12

AI Companies Destroy Books for Model Training as Google’s 2009 Scan Method Sits Unused

2 articles · Updated · Ars Technica · Aug 12

Summary

  • Physical books are being cut apart, scanned and discarded so AI companies can feed long-form texts into models at the lowest cost and highest speed.
  • That approach reflects a race for better training data: books offer engaging, high-quality prose, and destructive scanning is faster and cheaper than preserving each volume.
  • Google patented a non-destructive scanning system in 2009, but it is slower, costlier and imperfect—page curvature can distort text, pages can be missed, and rushed handling creates glitches.
  • Libraries and preservation groups such as the Internet Archive take the opposite approach, treating older books as materials that require careful, time-intensive human handling to avoid permanent loss.

Insights

Why did a massive billion-dollar piracy settlement suddenly make destroying millions of physical books the cheapest option for AI?
Could the AI industry's secret quest for legal data be permanently erasing the world's rarest physical books?