Recently unsealed court documents from the high-profile legal battle between The New York Times and tech giants OpenAI and Microsoft have exposed remarkably candid internal communications. The papers reveal that executives and researchers within both companies openly warned that their aggressive data-harvesting practices were triggering a destructive "doom loop," threatening the economic foundations of the web, and functioning as what one top Microsoft scientist characterized as the "largest theft of labor in human history."
The 92-page filing provides an unprecedented look behind the curtain of the generative AI boom, offering concrete proof that the architects of these foundational models were well aware of the legal, ethical, and economic fallout their products would inflict on creators and publishers. Despite internal red flags warning that large language models (LLMs) would cannibalize their own content supply chains and devastate referral traffic, both companies pushed forward relentlessly, motivated by the prospects of commercial dominance and massive financial windfalls.

Among the most striking disclosures in the filing are comments from Brent Hecht, Microsoft’s Director of Applied Science. In internal documentation cited by the plaintiffs, Hecht did not mince words regarding the scale of data harvesting required to power modern AI systems. He described the process as "an astonishing theft of unprecedented proportions" and warned that defending these practices under the umbrella of "fair use" would make "a complete mockery of the idea of fair use."
Microsoft has since attempted to distance itself from Hecht’s explosive assertions. Alex Haurek, a company spokesperson, told The Verge that the comments reflected one employee’s individual perspective, did not constitute a legal analysis, and did not represent the official views of the corporation. In a separate court filing, Jordan Usdan, General Manager for Data Strategy and Ops at Microsoft AI, further characterized Hecht’s role as adversarial, describing him as an academic thinker holding divergent and futuristic viewpoints who was not authorized to speak on behalf of Microsoft regarding the theoretical effects of AI on content creators.
Nevertheless, the unsealed documents show that Hecht’s fears—and those of various OpenAI employees—accurately predicted the current reality of the digital landscape, where AI-driven search overviews and conversational chatbots increasingly absorb web traffic that once went directly to publishers.

The comprehensive filing captures numerous provocative statements from prominent figures across both organizations, including Microsoft CEO Satya Nadella and OpenAI co-founder Sam Altman, painting a vivid picture of an industry knowingly disrupting the cultural and economic ecosystem it relies upon.
An Astonishing Theft and the Death of Fair Use
The legal arguments detailed in the court documents home in on the core tension of the lawsuit: the wholesale, unpermitted copying of millions of copyrighted articles to train commercial AI products that directly substitute for the original journalism. The filing highlights an admission from OpenAI’s Head of ChatGPT, who noted that publishers faced an "existential threat" from AI products that are "largely substitutive, period" and would only become more substitutive as the technology improved.

Legal experts note that such acknowledgments severely undermine the defendants’ reliance on the "fair use" doctrine. Because market substitution is a central vulnerability in copyright law—highlighted by recent Supreme Court precedent such as the Andy Warhol Foundation v. Goldsmith decision—these internal recognitions that chatbots replace the need to visit source material strike at the heart of the tech companies’ legal defense.
A "Doom Loop" Destroying the Content Supply Chain
The documents reveal that internal recognition of the damage was not limited to individual scientists. Satya Nadella acknowledged under oath that conversing with chatbots has substituted traditional search, providing information directly on the AI platform rather than requiring users to navigate to the underlying source website.

More damningly, an internal Microsoft document explicitly diagnosed the systemic problem, noting that the company’s AI content strategy had initiated a "doom loop." The memo warned that this loop would simultaneously hurt the performance of their own models and degrade the entire web: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’"
Further elaborating on this theme, Microsoft acknowledged that foundation models naturally compete with the very information work domains from which they draw their content. Because LLMs act as a substitute for the labor of the people who produced the original writing, newspapers, and books, executives recognized a "real risk" that generative AI would significantly disrupt the employment of the data generators, effectively creating a product that destroys its own supply chain.
Chasing "Gazillions" and Ignoring Paywalls

While corporate messaging often emphasizes the altruistic pursuit of technological advancement and global knowledge democratization, internal communications paint a starkly different financial motivation. Around the time these systems were scaling up, OpenAI co-founder Greg Brockman wrote candidly that he was "deeply motivated by the gazillions" he hoped to gain by commercializing the technology—a sentiment that aligns with subsequent reports of OpenAI planning massive public valuations.
This financial drive appears to have fostered a culture of looking the other way when it came to data acquisition ethics. Individuals within both OpenAI and Microsoft routinely bypassed paywalls and violated terms of service to amass training data. For instance, an OpenAI corporate representative testified during depositions that he was entirely unaware of any institutional effort to detect or remove paywalled content from their training datasets, despite Nadella later offering public platitudes suggesting that paywalled material ought to be properly licensed.
Regurgitation and the Loss of Referral Traffic

Internal OpenAI records also demonstrate acute awareness of technical hurdles regarding copyright compliance. As early as 2021, employees recognized that preventing model memorization was vital for minimizing copyright violations in generated outputs. Yet, by June 2022, team members acknowledged that GPT-4 had "memorized a ton of data and therefore will be insanely good at regurgitation." The plaintiffs’ filing lists numerous examples where ChatGPT responded to queries by outputting verbatim blocks of text straight from articles published by The New York Times, The Mercury News, The Denver Post, LifeHacker, and Eurogamer.
The impact of these capabilities on the broader publishing ecosystem is heavily quantified in the court papers through expert testimony. Dr. Goldfarb, an economic expert for OpenAI, found that declining referral traffic to The Times was driven by factors including Google AI Overviews, which he estimated may have depressed search referrals by 20 to 60 percent. Furthermore, media expert Dr. Sinnreich cited a 2026 Reuters Institute analysis showing that monthly referral traffic from Google Search and Google Discover dropped by billions of visits following the rollout of AI summaries.
As the litigation proceeds, Microsoft representatives have pushed back against interpretations of executive statements, with Haurek arguing that Nadella’s testimony addressed broad industry principles rather than specific legal conclusions regarding copyright. However, the newly unsealed records leave little ambiguity regarding what the companies suspected all along: that building powerful generative models on the uncompensated labor of the web would permanently alter, and potentially devastate, the creators of the culture they feed upon.
Leave a Reply