Skip to content
DIGITAL RIGHTS & COPYRIGHT LAW

Authors Sue OpenAI and Microsoft in New York Court Over Alleged Mass Copyright Infringement

Over the past three years, the legal landscape surrounding artificial intelligence has shifted dramatically as creators, publishers, and copyright holders push back against how tech companies gather data. Writers have filed a series of high-profile lawsuits accusing AI firms of training their massive language models on pirated books without permission or compensation. While some of these cases have already produced early rulings—including a bittersweet fair use victory secured by Meta in California—the legal battleground in New York has taken center stage.

In the U.S. District Court for the Southern District of New York, several separate lawsuits have been bundled into a single massive proceeding overseen by Judge Sidney Stein. This coordinated legal action includes the prestigious Authors Guild’s class action lawsuit, a separate case filed by a group of nonfiction writers who were the first to name Microsoft as a corporate defendant, and the ongoing Tremblay and Silverman lawsuit. The latter originally commenced in California in 2023, where it successfully survived a partial dismissal before being transferred to the New York docket.

This week, the plaintiffs in these consolidated cases reached a critical juncture by filing a motion for summary judgment. Ahead of any potential jury trial, the authors are asking Judge Stein to issue a definitive ruling that OpenAI copied their copyrighted works without authorization and that this unauthorized utilization cannot possibly qualify as fair use under United States copyright law. The newly filed motion specifically covers 194 distinct titles, asking the court for a formal finding of liability rather than immediate financial damages at this stage.

In their legal brief, the authors paint a stark picture of the modern publishing industry under pressure from rapidly advancing technology. "OpenAI’s GPT models pose an existential threat to those who write and publish books," the brief states, adding a warning about the current economic climate that "AI-generated books of all types are already flooding the market."

Built on Mass Piracy

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

The core of the authors’ legal argument begins with how OpenAI allegedly acquired the digital copies of books used during the formative stages of its model training. While portions of the newly submitted court filing remain heavily redacted to protect proprietary or sensitive information, the documents explicitly accuse OpenAI of utilizing unauthorized, torrented copies downloaded from notorious shadow libraries.

"OpenAI did not even buy the books it used. Instead, it began by torrenting [REDACTED] books from the notorious and illegal pirate library Library Genesis, also known as LibGen," the motion reads.

At the time OpenAI was allegedly sourcing these materials, Library Genesis had already earned a prominent spot on the U.S. Trade Representative’s official list of notorious piracy markets. According to the plaintiffs, OpenAI leadership and engineering teams were fully aware of the highly controversial and illegal nature of the platform when they acquired the datasets.

Furthermore, the motion alleges that OpenAI actively "took steps to conceal their piracy from the public." As evidence, the filing points to the academic paper originally published by OpenAI to introduce the GPT-3 model. In that documentation, OpenAI allegedly relabeled book compilations that internal drafts had previously identified as "Libgen1" and "Libgen 2," changing the nomenclature to the much more nondescript labels of "Books1" and "Books2."

"OpenAI employees understood at the time that they had sourced books from an illegal site," the court filing notes.

The obfuscation reportedly did not stop with a simple name change in research papers. According to the motion, OpenAI deliberately "deleted its LibGen files in the summer of 2022 due to legal concerns." The authors emphasize in their brief that these specific datasets represent "the only two training corpuses OpenAI has ever deleted."

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Before those files were purged from its systems, OpenAI allegedly leveraged the pirated library to train its early foundational models. Summarizing this phase of the corporate lifecycle, the authors write in their brief that the company essentially "built the foundations of its business on mass piracy."

Replacing George R.R. Martin

While the allegations surrounding torrenting and digital piracy form a major pillar of the lawsuit, the motion also tackles the broader economic motivations and long-term goals of the company. The plaintiffs allege that OpenAI intentionally built and refined its models with the explicit objective of replacing the human authors whose copyrighted works it had copied. To support this assertion, the legal team highlights controversial public statements made on social media by a key corporate employee.

In 2022, OpenAI brought on Tarun Gogineni to spearhead its efforts aimed at improving the literary and writing quality of its artificial intelligence models. According to the summary judgment motion, Gogineni was fully aware that the systems he was helping to train would ultimately displace professional authors, a development he allegedly viewed as an "acceptable economic disruption."

This aspect of the case carries particular weight because Gogineni specifically referenced one of the named plaintiffs in the litigation: acclaimed author George R.R. Martin, best known for penning the epic fantasy series A Song of Ice and Fire, which served as the basis for the hit HBO television adaptation Game of Thrones.

Nearly two years after Martin joined the lawsuit against the company, Gogineni took to social media in 2025 to outline what he described as his "research mission." He publicly stated that his goal was to have GPT models write the "last two books of [Martin’s] A Song of Ice and Fire."

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

"Even if [Martin] dies early, GPT-5 will autocomplete his series," Gogineni added in the post, a statement that the plaintiffs’ legal team highlights as clear evidence of the company’s intent to substitute human creative labor with automated generation.

Not Fair Use

The question of whether training artificial intelligence models on copyrighted text constitutes lawful fair use remains the central legal debate across all modern AI copyright litigation. OpenAI and other major tech firms have consistently maintained that ingesting vast corpuses of text to train models falls safely within the boundaries of fair use. While federal courts have shown some willingness to entertain aspects of this defense in recent rulings, judges have also established crucial caveats that complicate the tech industry’s legal position.

In their motion, the authors cite the recent 2025 ruling in Bartz v. Anthropic, a California case that established that while training an AI model on text could theoretically be considered fair use under certain circumstances, downloading those texts from an unauthorized pirate library certainly is not. That court ruled that pirating books that are readily available for purchase through legal commercial channels is "inherently, irredeemably infringing."

Building on this precedent, the New York plaintiffs argue that the copying of their specific literary works was entirely avoidable for general-purpose model training, and that possessing their copyrighted books was never strictly necessary to create a generalized AI system.

Broader Claims Against Microsoft

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

The legal action in New York is not exclusively directed at OpenAI. The newly filed summary judgment motion also asks Judge Stein to hold Microsoft vicariously liable for OpenAI’s alleged copyright infringement. The plaintiffs argue that Microsoft maintained the legal authority and practical ability to supervise OpenAI’s corporate conduct while standing to reap massive financial benefits from the enterprise.

To illustrate the depth of this commercial partnership, the authors emphasize that Microsoft has invested roughly $13 billion into OpenAI through three separate major financial agreements signed in 2019, 2021, and 2023.

OpenAI has not yet responded directly to the specific allegations raised in this latest round of filings, but the company has made its overarching legal stance clear in parallel proceedings. In a cross-motion for summary judgment filed on the same day, OpenAI argued that its utilization of copyrighted books constitutes fair use as a matter of law, and further asserted that any instances of its models regurgitating copyrighted text verbatim are vanishingly rare.

The filings submitted this week represent only one front in a much broader, escalating legal campaign against generative AI developers. Over recent days, other major media and publishing plaintiffs—including The New York Times, the Daily News, and the Center for Investigative Reporting—submitted a combined summary judgment motion of their own against both OpenAI and Microsoft.

With many millions of dollars in potential damages at stake, alongside the foundational future of how artificial intelligence models are trained and commercialized, these high-stakes intellectual property cases will undoubtedly be contested aggressively by all parties involved.

A copy of the authors’ redacted motion for partial summary judgment has been filed at the U.S. District Court for the Southern District of New York.

Leave a Reply

Your email address will not be published. Required fields are marked *