Skip to content
DIGITAL RIGHTS & COPYRIGHT LAW

Sony and Warner Sue Anthropic Over Massive Pirate Data Haul Used in AI Training

Major music publishers, including industry giants Sony and Warner, have launched a fresh legal assault against artificial intelligence company Anthropic in the U.S. District Court for the Northern District of California. The late-Friday complaint alleges that a massive data haul previously used by the AI firm to train its models included thousands of copyrighted songbooks, sheet music collections, and lyrics.

This new filing follows a massive $1.5 billion settlement reached last September between Anthropic and a class of book authors over the unauthorized use of roughly seven million pirated titles. While that multi-billion-dollar resolution closed one major chapter of litigation, it failed to deter other rightsholders from stepping forward to pursue their own independent legal claims over the same underlying data acquisition methods.

The newly minted lawsuit argues that the digital hoarding operation went far beyond standard web scraping, involving the direct downloading and torrenting of copyrighted materials through shadow libraries. According to the legal complaint, each pirated work torrented by the defendants was likely shared thousands, if not tens of thousands, of times across the peer-to-peer network, thereby depriving music publishers of substantial revenue.

The publishers’ complaint brings forward allegations of direct and contributory copyright infringement stemming from the active torrenting process itself, alongside broader infringement claims tied to subsequent AI training practices. Furthermore, the lawsuit cites Digital Millennium Copyright Act (DMCA) violations for the systematic stripping of copyright management information and notices. Notably, the legal action names Anthropic CEO Dario Amodei and co-founder Benjamin Mann personally as individual defendants alongside the corporation.

“A Cute Little LibGen Babysitter”: Music Publishers Sue Anthropic Founders Over Torrenting Spree

“A Cute Little LibGen Babysitter”

The foundations of the music publishers’ case rely heavily on established facts, internal communications, and evidentiary admissions unearthed during earlier Anthropic litigation. Much of the quoted material and technical context is drawn directly from the Bartz v. Anthropic case, the very legal proceeding that ultimately forced the company into its historic $1.5 billion settlement with aggrieved book authors.

According to disclosures from the Bartz proceedings, co-founder Benjamin Mann discussed the active torrenting of materials from Library Genesis (LibGen) quite openly inside Anthropic’s internal Slack channels. At one point, Mann even shared a screenshot of his ongoing downloading activity with his engineering colleagues. He famously described a custom program he had written to manage and automate the massive download process as “a cute little libgen babysitter,” as detailed in the court filing.

The complaint further underscores that Anthropic’s leadership was acutely aware of the dubious legal standing of the platforms they were utilizing. Internal messaging records reveal that Mann characterized LibGen itself as being “sketchy AF.” Other internal communications show that members of Anthropic’s Archive Team went even further, openly designating the operations as a “blatant violation of copyright.”

Despite these internal warnings and acknowledgments regarding the legality of the source material, CEO Dario Amodei ultimately approved the torrenting initiative. The complaint notes that Dr. Amodei explicitly admitted that Anthropic had many legitimate commercial avenues from which it could have legally purchased these copyrighted works for model training, but instead chose to torrent them because executing the downloads through shadow libraries was significantly faster and entirely free.

“A Cute Little LibGen Babysitter”: Music Publishers Sue Anthropic Founders Over Torrenting Spree

“A Popular (and Illegal) Library”

As the data acquisition efforts expanded, internal enthusiasm among engineering staff remained high. When Mann discovered that the Pirate Library Mirror (PiLiMi) was primed and ready for mass torrenting in the summer of 2022, he circulated the link directly to his colleagues with the comment, “just in time!” Another Anthropic employee enthusiastically responded to the update by proclaiming, “zlibrary my beloved,” according to excerpts included in the legal filing.

Following this discovery, Anthropic engineers systematically compared the five million books they had already illegally downloaded from LibGen against the roughly seven million titles available via PiLiMi. They subsequently proceeded to download the remaining two million unique items that were missing from their initial haul. Internal records demonstrate that employees maintained a clear understanding of the nature of these repositories, describing PiLiMi internally as “a popular (and illegal) library.”

The music publishers’ complaint alleges that this aggressive torrenting sweep inadvertently netted hundreds of specific songbooks, vocal collections, and instrumental sheet music editions. Exhibit A attached to the legal filing details a wide array of specific titles targeted in the haul, including prominent works such as The Beatles Complete Scores, the Best of Taylor Swift Songbook, and Bon Jovi’s These Days.

Even as internal sentiment within the company evolved and staff became somewhat hesitant about continuing to train frontier AI models on pirated material due to escalating legal risks, the company allegedly chose to retain the acquired files within its central corporate library archives rather than purging them.

“A Cute Little LibGen Babysitter”: Music Publishers Sue Anthropic Founders Over Torrenting Spree

Rewriting Pirate Library History

While the core allegations concerning Anthropic’s torrenting activities are heavily anchored in well-documented court records and internal chat logs, the complaint’s narrative regarding the broader history of shadow libraries contains noticeable historical inaccuracies. The legal filing asserts, for instance, that the Federal Bureau of Investigation shut down LibGen in late 2021, prompting pirates to subsequently copy its database contents to establish Z-Library.

In reality, LibGen was never successfully shut down by federal authorities and remains actively online today. Conversely, Z-Library was originally founded much earlier, in 2008, initially operating as a LibGen mirror before eventually expanding into one of the largest independent pirate ebook repositories on the internet. It was actually Z-Library—not LibGen—that lost its primary domains in a sweeping international law enforcement action led by the FBI in November 2022, several months after Anthropic had already wrapped up its major data downloading sprees.

While these historical missteps in the complaint do not alter the fundamental legal accusations regarding unauthorized acquisition and usage, they highlight a surprisingly shaky grasp of shadow library chronology within a lawsuit that is otherwise built meticulously upon the operational details of digital piracy.

One Torrenting Spree, Three Lawsuits

The initial two counts outlined in the complaint specifically target the underlying peer-to-peer torrenting activity itself, independent of the subsequent AI training phases. Because the BitTorrent protocol inherently operates by simultaneously uploading data fragments to other users while a client is downloading, the publishers argue that Anthropic did not merely reproduce their copyrighted works internally; it actively participated in distributing them to countless unknown third parties across the global network.

“A Cute Little LibGen Babysitter”: Music Publishers Sue Anthropic Founders Over Torrenting Spree

This legal theory is hardly novel, having been deployed by corporate rightsholders against individual peer-to-peer file-sharing users for over two decades. However, its application here targets a major artificial intelligence corporation reportedly eyeing a massive multi-trillion-dollar valuation, which the publishers claim effectively sustained and normalized the broader BitTorrent piracy ecosystem through its corporate data collection practices.

This latest legal filing represents the third major lawsuit to emerge directly from Anthropic’s historical torrenting spree. Following the book authors’ landmark $1.5 billion settlement and a parallel lawsuit filed in January by a group of music publishers that included Concord, Universal, and BMG, industry observers note that this new complaint may still not be the last of its kind.

Anthropic maintains a firmly opposing stance, arguing that its data acquisition and model training activities are fully protected under the legal doctrine of fair use. A company spokesperson told Ars Technica that the new filing constitutes the third consecutive lawsuit brought forward by the same legal representation, recycling allegations that are already undergoing formal review in existing court proceedings. The spokesperson emphasized that AI training constitutes fair use, as previously affirmed by the court in the Bartz litigation, and confirmed that the company intends to vigorously defend itself against the claims.

Legal analysts point out, however, that the previous fair use rulings applied specifically to the downstream training phase of the AI models rather than the initial acquisition phase. Indeed, the presiding court previously observed that the underlying act of downloading millions of books via torrent networks amounted to "straightforward piracy but at massive scale."

“A Cute Little LibGen Babysitter”: Music Publishers Sue Anthropic Founders Over Torrenting Spree

In their prayer for relief, the music publishers are demanding up to $150,000 in statutory damages for each individual infringed work. Given that the lists of affected titles encompass tens of thousands of specific compositions and songbooks, potential financial exposure could stretch into the billions of dollars if the publishers prevail in court.

Leave a Reply

Your email address will not be published. Required fields are marked *