Excerpts from internal communications, made public through legal filings in the New York Times' lawsuit against Microsoft and OpenAI, shed new light on how executives at both companies privately viewed the risks tied to training their models on copyrighted material. According to briefs cited by Tom's Hardware, a Microsoft director reportedly described the scraping of data for AI training as the largest theft of labor in human history, a striking admission coming from within a company itself named as a defendant.
On OpenAI's side, an executive is said to have acknowledged that ChatGPT posed an existential threat to publishers, a statement that could carry significant weight in assessing the harm alleged by the New York Times. The newspaper accuses both companies of using millions of its articles without permission or compensation to train large language models, and of allowing those models to reproduce or summarize its content in ways that directly compete with its own journalism.
The disclosures come as several similar lawsuits move through U.S. courts, pitting publishers, authors and artists against leading generative AI developers. The central legal question remains whether training on copyrighted material qualifies as fair use or amounts to outright infringement, an issue American courts have yet to resolve consistently across cases.
If these internal quotes are confirmed and deemed admissible, they could strengthen the New York Times' position by showing that executives at both companies were aware of the potential harms before rolling out their products. Closely watched across the media industry, the case could become a defining precedent for the broader generative AI sector.