Anthropic’s New Music Lawsuit Is Really a Data-Lineage Test
Sony and Warner’s publishers allege mass piracy, lyric copying, and stripped attribution; the decisive issue is not whether every model training run is illegal, but whether AI companies can prove where training works came from and what their systems reproduce.
By OMIKINA Editorial · Published · Updated through
Key points
- Music publishers linked to Sony and Warner allege that Anthropic acquired copyrighted compositions through torrents and scraping, used them in training, reproduced lyrics in outputs, and removed copyright-management information. These are allegations, not findings. Sources: S1, S2
- Anthropic disputes the claims and says it will defend itself. The company’s response must appear beside the publishers’ allegations because the case has only begun. Sources: S3
- The U.S. Copyright Office says fair-use analysis for generative-AI training depends on the facts, including the works used, how they were obtained, the purpose of the use, safeguards, and effects on markets. Sources: S4
- A prior federal order in the Bartz books case treated Anthropic’s training use as fair use on that record but did not excuse building a central library from pirated copies. That order does not decide this new music case. Sources: S5
The complaint makes four claims, not one
The publishers’ complaint alleges that Anthropic obtained copyrighted musical compositions through torrents and web scraping, copied those works into training systems, produced protected lyrics in some Claude outputs, and removed or changed copyright-management information. The suit concerns compositions and lyrics. It should not be described as a case about ownership of every sound recording or a performer’s voice.
None of those allegations has been proven. A complaint states the plaintiffs’ case before evidence is tested. Anthropic says it disagrees and will defend itself. The publishers will still need to show ownership, copying, legally actionable conduct, and the basis for any claimed damages.
Training use and data acquisition are different questions
A court can ask whether training on a work was fair use and separately ask whether the copy used for training was lawfully obtained. That distinction matters here because the publishers do not only challenge model training. They allege that Anthropic first built or used collections containing pirated copies.
The Copyright Office’s training report rejects a single answer for every system. It says the analysis depends on facts such as the type of work, the purpose of the use, how a copy was acquired, what safeguards were used, and whether the system harms a market for the original or a licensed use. A ruling about one dataset or model may not control another.
The Bartz order shows why provenance matters
In the earlier Bartz books case, Judge William Alsup ruled that Anthropic’s training use was fair use on the evidence before him. The same order said that using pirated books to build a permanent central library was not justified by the later training purpose. The case later settled, so that mixed order did not become a final national rule for all AI training.
The music publishers are using a similar separation in their new complaint: acquisition, training, output, and attribution are presented as distinct acts. The Bartz order may shape the arguments, but it does not prove that the new allegations are true or decide whether lyrics, licensing markets, and Claude outputs create a different result.
The operational test is a traceable data chain
AI companies now need records that connect a training item to its source, acquisition method, license or legal basis, retention policy, and model version. They also need output tests for close reproduction and controls that preserve copyright-management information when the law requires it. Those records cannot decide fair use by themselves, but they make a company’s account testable.
The next evidence will come from Anthropic’s formal response, discovery about the datasets, examples of disputed outputs, ownership records, and court rulings on each claim. Until then, the accurate conclusion is narrow: major publishers have raised a serious provenance challenge, Anthropic denies liability, and no court has decided this case.
Why it matters
Copyright risk now reaches beyond a general debate about whether training can be fair use. It reaches the full data chain: where a work came from, what rights or legal basis covered the copy, what the model retained, what it can reproduce, and whether attribution survived. The case could make provenance and output testing core infrastructure for AI companies, even before a final ruling.
Sources
- Complaint, Sony Music Publishing US LLC et al. v. Anthropic PBC et al. — United States District Court for the Northern District of California ·
- Sony Music Publishing US LLC et al. v. Anthropic PBC et al., docket 5:26-cv-09217 — Justia Dockets ·
- Anthropic hit with major lawsuit from music publishers — Axios ·
- Copyright and Artificial Intelligence, Part 3: Generative AI Training — United States Copyright Office ·
- Order on Fair Use, Bartz v. Anthropic PBC — United States District Court for the Northern District of California ·
Read OMIKINA's editorial standards · Review corrections · Follow the AI-narrated podcast · Follow the RSS briefing