LIVE
Court filings reveal internal admissions on AI content scraping19/09/26|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Court filings reveal internal admissions on AI content scraping19/09/26|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|
Regulation

Court filings reveal internal admissions on AI content scraping

Legal briefs filed in the New York Times lawsuit against Microsoft and OpenAI quote a Microsoft director calling AI data scraping the largest theft of labor in history, and an OpenAI executive describing ChatGPT as an existential threat to publishers.

September 19, 20263 min readPublished byHacker News

Excerpts from internal communications, made public through legal filings in the New York Times' lawsuit against Microsoft and OpenAI, shed new light on how executives at both companies privately viewed the risks tied to training their models on copyrighted material. According to briefs cited by Tom's Hardware, a Microsoft director reportedly described the scraping of data for AI training as the largest theft of labor in human history, a striking admission coming from within a company itself named as a defendant.

On OpenAI's side, an executive is said to have acknowledged that ChatGPT posed an existential threat to publishers, a statement that could carry significant weight in assessing the harm alleged by the New York Times. The newspaper accuses both companies of using millions of its articles without permission or compensation to train large language models, and of allowing those models to reproduce or summarize its content in ways that directly compete with its own journalism.

The disclosures come as several similar lawsuits move through U.S. courts, pitting publishers, authors and artists against leading generative AI developers. The central legal question remains whether training on copyrighted material qualifies as fair use or amounts to outright infringement, an issue American courts have yet to resolve consistently across cases.

If these internal quotes are confirmed and deemed admissible, they could strengthen the New York Times' position by showing that executives at both companies were aware of the potential harms before rolling out their products. Closely watched across the media industry, the case could become a defining precedent for the broader generative AI sector.

Tags
copyrightlawsuitopenaimicrosoftjournalismtraining-data

Read also

Court filings reveal internal admissions on AI content scraping · nAIvigate