According to newly unredacted court filings cited by TechCrunch, a Microsoft executive reportedly made internal remarks sharply critical of the data collection practices used to train large language models, going as far as describing large-scale web scraping as the largest theft of labor in human history. These private comments stand in stark contrast to the company's public messaging, which has consistently defended the use of widely available web data to train its AI systems, including through its partnership with OpenAI.
The disclosure emerges from a broader legal battle involving publishers, authors, and content creators who accuse technology companies of using copyrighted material without permission or compensation to train their models. Microsoft, like several other major tech firms, faces multiple lawsuits over the use of copyrighted content, including news articles and literary works, within its AI training pipelines.
The fact that these remarks came from within the company itself gives the disclosure particular weight. It suggests that internal doubts existed, even among executives involved in AI strategy, about the legitimacy of certain large-scale data collection practices. Documents obtained through court unsealing are often leveraged by plaintiffs to demonstrate prior awareness of potentially problematic conduct.
The episode highlights the growing tension between the rapid advance of generative AI and unresolved questions around intellectual property. As U.S. courts begin issuing rulings on how fair use doctrine applies to model training, this kind of internal admission could influence ongoing settlement negotiations between tech companies and rights holders, as well as the outcome of several pending lawsuits.