Skip to content
TORRENT & P2P MEDIA NEWS

U.S. Magistrate Judge Shields OpenAI and Anthropic from Subpoenas in High-Stakes Meta Copyright Lawsuit

Over the past two years, copyright holders across the literary, academic, and publishing sectors have increasingly turned to the federal courts, filing a sweeping wave of lawsuits against the developers of advanced artificial intelligence models. These legal challenges target the fundamental mechanics of how large language models are built, specifically focusing on the datasets used to train them. Among the primary targets of this ongoing litigation is Meta Platforms, which currently finds itself defending against a lengthy roster of copyright infringement claims.

One of the most prominent legal hurdles facing the social media giant is a class-action lawsuit originally brought forward by notable authors, including Richard Kadrey and Sarah Silverman. This foundational complaint accused Meta of training its flagship Llama family of models using unauthorized, pirated digital books. Furthermore, the authors alleged that the company actively participated in peer-to-peer distribution networks by sharing those copyrighted files with other users on the BitTorrent network while downloading the training materials.

The Evolution of Meta’s Fair Use Defense

Last summer, U.S. District Judge Vince Chhabria delivered a nuanced ruling in the case, determining that the actual training of artificial intelligence models using copyrighted texts constituted fair use under U.S. copyright law. However, that bittersweet victory for Meta left a critical loose end. The claims centered around unauthorized BitTorrent distribution remained active and alive, leaving the company vulnerable to continued litigation regarding its download and sharing methods.

Earlier this year, Meta introduced an entirely new line of defense to address these lingering peer-to-peer distribution claims. In a supplemental interrogatory response filed with the court, the tech conglomerate argued that any uploading of pirated books that occurred automatically during its torrent downloads was merely "part-and-parcel" of pursuing an otherwise lawful fair use purpose.

Meta elaborated on this argument by asserting that BitTorrent represented a significantly more efficient and reliable means of obtaining massive training datasets. In the specific case of Anna’s Archive, an underground shadow library, the company maintained that torrenting was practically the only viable way to acquire the data in bulk quantities. Because the architecture of the BitTorrent protocol inherently relies on users uploading data to one another while downloading, Meta contended that any subsequent sharing of files was simply an unavoidable, inherent characteristic of the technology itself.

Meta’s novel torrenting defense is currently being thoroughly tested across three related lawsuits filed in the same jurisdiction. These coordinated actions were brought by Chicken Soup for the Soul, academic publisher Cognella, and Cambronne Inc., an entity represented by journalist John Carreyrou. All three of these related cases have been assigned to Judge Chhabria and focus squarely on the same shadow library torrenting activities that have drawn intense judicial scrutiny.

Publishers Seek Evidence From AI Rivals

Rather than waiting for Meta to fully document and disclose the proprietary technical details surrounding its torrent client setup and network configurations, the plaintiff publishers decided to pursue an aggressive alternative discovery strategy. They looked outward, targeting two of Meta’s primary artificial intelligence competitors who might hold valuable industry-wide insight into how large language models acquire shadow library data.

Rightsholders Can’t Use OpenAI and Anthropic to Dismantle Meta’s Seeding Defense

In August, the publishers issued formal subpoenas to both OpenAI and Anthropic. The legal demands required the two AI companies to hand over the identities, specific software versions, and detailed configurations of every torrent client they had utilized since 2019. Crucially, these subpoenas sought any existing records documenting internal efforts or technical configurations designed to prevent the uploading or seeding of files during the download process.

In the context of other similar copyright lawsuits winding through the federal court system, both OpenAI and Anthropic have previously admitted to utilizing books sourced from shadow libraries to train their own language models. The publishers reasoned that if OpenAI or Anthropic had successfully configured their respective torrent clients to suppress uploading while downloading, Meta’s core legal argument regarding technical necessity would be severely undermined.

"If OpenAI torrented but configured its clients to suppress uploading, then the redistribution Meta calls an ‘inherent characteristic’ of the protocol was a setting Meta declined to change," the publishers argued in court filings. This strategic line of argumentation directly builds upon an earlier revelation in the legal proceedings, which uncovered that a Meta engineer had previously written a custom script specifically designed to prevent seeding, while apparently leaving the leeching functionality untouched.

OpenAI and Anthropic Push Back Against Subpoenas

Rather than complying with the broad document demands, OpenAI and Anthropic strongly pushed back against the subpoenas, offering resistance in the form of formal legal filings. The AI companies informed the court that examining the technical features and operational logs of their own torrent clients would yield little to no relevant evidence regarding Meta’s specific practices. They argued that corporate practices at OpenAI or Anthropic had no bearing on what Meta chose to do behind closed doors.

"Clients are not interchangeable, they differ in their default upload settings, in whether those defaults can be reconfigured, and in their capacity to suppress uploading during and after a download," lawyers representing Anthropic wrote in their opposition papers. "What Anthropic’s client allowed shows nothing about what Meta’s did."

OpenAI echoed these exact sentiments in its own communications with the court, pointing out that the plaintiffs had presented zero evidence demonstrating that OpenAI had employed the exact same torrent clients or successfully built comparable corpora of training data to the massive datasets amassed by Meta.

Ahead of the formal judicial ruling, the publishers attempted a compromise, offering to drop the expansive document demands entirely if either OpenAI or Anthropic would simply explain how they acquired their shadow-library data and clarify whether they made deliberate attempts to prevent uploading. When the competing AI firms rejected this compromise, the matter was left in the hands of the magistrate judge to decide.

Rightsholders Can’t Use OpenAI and Anthropic to Dismantle Meta’s Seeding Defense

Magistrate Judge Sides With AI Competitors

In a definitive order released last week, U.S. Magistrate Judge Thomas Hixson formally sided with OpenAI and Anthropic, quashing the subpoenas and shielding the competing AI companies from having to turn over their proprietary torrent logs. Without rendering a final judgment on the ultimate validity of Meta’s broader seeding arguments, the court concluded that the internal torrent logs of rival AI developers were simply not the appropriate or most reliable source for gathering this type of evidence.

"To the extent Meta’s fair use defense hinges on the assertion that its use of BitTorrent was the only way BitTorrent can be used, that assertion can be tested by examining the BitTorrent client itself," Judge Hixson wrote in his ruling.

The magistrate judge emphasized that asking OpenAI or Anthropic for their historical logs provides virtually no insight into Meta’s distinct technical setup or the inherent technical capabilities of the software clients involved. Furthermore, Judge Hixson noted that because any standard user of a torrent client could theoretically be relevant under that expansive definition, the plaintiffs’ legal logic lacked practical boundaries. He questioned why the plaintiffs’ own technical experts could not simply utilize the torrent clients independently to demonstrate how such software can be operated.

Regarding the publishers’ assertions that shadow library data could only be downloaded in bulk through torrent protocols, Judge Hixson pointed out that this premise could easily be verified by communicating directly with the shadow libraries themselves, rather than attempting to compel third-party competitors to surrender sensitive operational data.

Focus Shifts to Meta’s Own Server Data

While the court ultimately decided that dragging competing AI developers into the discovery phase was entirely off-limits, a radically different standard was applied to data originating directly from Meta’s own internal corporate servers.

On September 11, Judge Hixson granted a significant motion in the related class-action lawsuit filed by Richard Kadrey and his fellow authors. This specific judicial order commanded Meta to surrender the command history files for every single server the tech giant used to execute its torrenting operations, which explicitly encompasses its internal virtual machines and Amazon Web Services instances.

In technical environments, command history files represent the persistent logs kept by a server recording every operational command typed by an active system operator. In the context of a machine dedicated to torrenting, these records presumably detail how the chosen torrent client was initially installed, alongside any modifications, tweaks, or scripts applied to its default upload settings.

Rightsholders Can’t Use OpenAI and Anthropic to Dismantle Meta’s Seeding Defense

This discovery order traces back to earlier in the year when Meta was forced to admit that it had improperly withheld relevant documents until well after the formal discovery deadline had already passed. To remedy this procedural failure, Judge Chhabria granted the plaintiffs extended discovery privileges, specifically unlocking records that illustrate how Meta’s torrent clients were configured and operated in practice.

Although Meta argued strenuously that the log files it had already voluntarily surrendered were more than sufficient to address the plaintiffs’ inquiries, Judge Hixson firmly disagreed. He ordered the company to hand over the comprehensive command histories without further delay.

The plaintiff authors remain hopeful that these detailed server command histories will finally unveil the precise identities of the copyrighted works that Meta downloaded via torrent networks. Whether this forensic data will definitively prove unauthorized seeding remains to be seen as the litigation progresses.

For the time being, the central question of whether Meta could have successfully downloaded the contested books without simultaneously seeding them back to the network rests squarely in the hands of the plaintiffs’ technical experts. Those experts are expected to utilize Meta’s newly acquired server records to formulate their positions, with their opening expert reports in the ongoing Meta cases due later this month.

Leave a Reply

Your email address will not be published. Required fields are marked *