Unsealed filings in the NYT copyright suit show Microsoft and OpenAI executives privately called AI training a form of theft
Newly unredacted filings in The New York Times' copyright lawsuit against Microsoft and OpenAI, made public September 17, 2026, show senior executives at both companies privately describing their AI training practices in far starker terms than either company has used publicly.
What's new
The filings quote a January 2024 internal memo from Brent Hecht, Microsoft's director of applied science, who wrote that mass ingestion of copyrighted publisher content for AI training was "an astonishing theft of unprecedented proportions" and, in his view, potentially "the largest theft of labor in human history." Hecht also wrote that "it is highly unusual that an end-product threatens the economic foundations of its essential suppliers" — a reference to AI chatbots undercutting the publishers whose journalism trained them.
On the OpenAI side, the filings cite Nick Turley, the head of ChatGPT, describing the threat to publishers as "existential" because chatbot answers are "largely substitutive" for the original reporting. Microsoft CEO Satya Nadella is quoted from deposition testimony saying "anything that is paywalled should be licensed by anyone who wants to use it ... for grounding or training," a position at odds with how the Times alleges its own paywalled content was actually used. The unsealed material also details the scale the Times says was involved: OpenAI's mid-training datasets allegedly contain more than 91,692 copies of articles from the Times, the Daily News, and the Center for Investigative Reporting, and a Common Crawl-derived dataset allegedly included more than 2 million documents from nytimes.com alone.
Context
The New York Times sued OpenAI and Microsoft in late 2023, alleging that ChatGPT and Copilot were trained on millions of Times articles without permission or payment, and that the models could reproduce substantial portions of that reporting verbatim. The case has since become one of the highest-profile tests of whether training generative AI models on copyrighted journalism qualifies as fair use. These filings were submitted as part of the Times' push for summary judgment, and unsealing them exposes internal admissions that neither company had disclosed publicly.
Why it matters
The gap between how AI companies describe their own training data practices internally and how they defend those practices in public and in court is now part of the public record in one of the most closely watched AI copyright cases. If a court credits these statements as evidence that the companies understood their scraping practices as harmful to publishers, it could weaken the fair-use defense both companies are relying on — a defense that most other generative-AI copyright suits, across text, image, and audio models, are also counting on holding up.
Corroborating sources
- Washingtonpost
https://www.washingtonpost.com/business/2026/09/17/microsoft-exec-called-ai-largest-theft-labor-history-court-records-show/
- Techcrunch
https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/
“New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.”