Updated
Updated · The Verge · Sep 4
Microsoft Says 8.2 Million Copilot Logs Show Under 1% News Text Reproduction
Updated
Updated · The Verge · Sep 4

Microsoft Says 8.2 Million Copilot Logs Show Under 1% News Text Reproduction

1 articles · Updated · The Verge · Sep 4

Summary

  • Microsoft told the court that fewer than 1% of 8.2 million Copilot chat logs contained at least 16 words matching news content, arguing the chatbot rarely reproduces material that could replace original articles or books.
  • 59,545 conversations in that dataset showed 16-word overlaps with grounded news content, Microsoft said, while an expert in the authors' case found only 24 responses with at least 30 matching words across 212 books.
  • Microsoft says those results support its fair-use defense, arguing copyrighted works were used to train systems for a different, transformative purpose even if some passages are occasionally reproduced.
  • The filing comes as Microsoft seeks summary judgment in consolidated lawsuits from The New York Times, the Center for Investigative Reporting and authors, who say Microsoft and OpenAI trained competing products on their works.

Insights

If Microsoft's AI rarely copies text verbatim, does summarizing a news article still steal the publisher's economic value?
Could a court ruling against Microsoft's fair use defense force the entire generative AI industry to rewrite its algorithms?
With built-in copyright detectors in enterprise AI, is Microsoft quietly admitting consumer AI models are legally vulnerable?