Updated
Updated · TechCrunch · Sep 17
NYT Filings Show OpenAI, Microsoft Logged 2 Million Times Articles as Executives Called Scraping Theft
Updated
Updated · TechCrunch · Sep 17

NYT Filings Show OpenAI, Microsoft Logged 2 Million Times Articles as Executives Called Scraping Theft

3 articles · Updated · TechCrunch · Sep 17

Summary

  • Unsealed filings in The New York Times copyright case show Microsoft and OpenAI executives privately described AI scraping as “theft” and publisher-facing chatbots as an “existential threat,” undercutting their fair-use defense.
  • Microsoft’s own data said Copilot cut click-throughs to nytimes.com by as much as 93% versus traditional Bing search, while internal documents warned of a “doom loop” that could damage both publishers and the web.
  • The filings allege the companies built training sets through mass scraping, Bing-index transfers and projects including Mango and Taxi, with one Common Crawl-derived dataset containing more than 2 million nytimes.com documents.
  • OpenAI’s mid-training datasets alone allegedly held 91,692 copies of works from the NYT, Daily News and Center for Investigative Reporting, and employees discussed bypassing the NYT paywall and stripping copyright notices from data.
  • The new material escalates a three-year lawsuit over whether training on copyrighted works is lawful, a question courts have often viewed favorably for AI firms even as the Trump administration recently backed OpenAI’s unlicensed training use.

Insights

Could explosive internal emails proving tech giants knowingly bypassed paywalls finally shatter the legal shield protecting generative AI?
If AI models destroy the news sources they learn from, what happens when they run out of human journalists to scrape?
With courts divided in 2026, will the revelation of a massive traffic drop force a total rewrite of global copyright laws?

The $300 Million Copyright Clash: How AI Training on 10.8 Million News Articles Sparked a Global Legal and Economic Crisis

Overview

The U.S. Department of Justice’s support for OpenAI and Microsoft in The New York Times copyright lawsuit marks a turning point in the battle over AI and news content. The DOJ argues that training AI on copyrighted material is fair use and vital for national security, but this stance has sparked backlash from publishers who face steep drops in web traffic due to AI answer engines. As AI models summarize and replace paywalled articles, publishers lose revenue and credibility, while tech companies sign multi-million dollar licensing deals with some media outlets. This legal and economic struggle is reshaping the future of journalism and copyright in the AI era.

...