Lawsuit tracker
U.S. District Court for the Northern District of California (Judge Yvonne Gonzalez Rogers) · ActiveAuthors v. Apple
Two novelists say Apple trained a language model on a dataset of pirated books — a claim showing that even a modest research model, not just a headline chatbot, can draw a class action years after the training run.
Key facts
- Filed September 5, 2025 in N.D. Cal. by authors Grady Hendrix and Jennifer Roberson against Apple.
- Alleges Apple trained its OpenELM models on Books3, a shadow-library dataset of roughly 200,000 pirated books.
- Consolidated in November 2025 with two related author suits, Martinez-Conde and Alexander; interim co-lead class counsel appointed January 2026.
- A consolidated class-action complaint, adding authors John Hornor Jacobs and Eboni McKinnon, was filed February 13, 2026.
- Apple answered on April 30, 2026 rather than moving to dismiss, denying the claims and asserting fair use.
What the complaint alleges
Grady Hendrix and Jennifer Roberson’s suit targets a specific, named dataset: Books3, a collection of nearly 200,000 books assembled from shadow libraries and widely used across the industry to train language models before its copyright problems became public. The authors say Apple used Books3 to train its OpenELM family of small, efficient language models, and separately allege that shadow-library books fed the foundation models behind Apple Intelligence, Apple’s on-device and cloud AI features, without disclosing which datasets were used.
Why a research model still counts
OpenELM is a research-oriented model family, not Apple’s flagship consumer product, and that is part of the point of this case: infringement claims do not require a chatbot with millions of users. If a dataset was pirated, using it to train any model — research or commercial, large or small — creates the same copying exposure. The suit argues that Apple’s own research publications about OpenELM’s training data gave plaintiffs a factual basis to allege Books3’s use directly, rather than inferring it from outputs.
Consolidation and early posture
The court consolidated Hendrix with two similar author actions, Martinez-Conde and Alexander, in November 2025, and appointed interim co-lead class counsel from Keller Rohrback and Susman Godfrey in January 2026. Plaintiffs then filed a consolidated class-action complaint in February 2026, adding authors John Hornor Jacobs and Eboni McKinnon to broaden the proposed class.
Apple answers instead of moving to dismiss
Rather than challenge the complaint’s legal sufficiency with a motion to dismiss, Apple filed an answer on April 30, 2026, denying the allegations and asserting fair use as an affirmative defense. That procedural choice matters: answering means Apple is prepared to litigate the fair-use question on the merits and through discovery, rather than trying to end the case early on the pleadings. It puts the dispute on a track similar to the merits-stage fights already underway in the OpenAI and Meta book cases.
Why it matters for training-data licensing
Books3 has already surfaced in multiple AI copyright suits, but Authors v. Apple shows its legal exposure is not limited to the largest model builders — any company that touched the dataset, for any purpose, inherits the same claims. For buyers, that means checking not just a vendor’s current data practices but the historical training runs behind any model in a supply chain, including smaller or discontinued ones. Taking a pirated dataset offline, as has happened with Books3, does not retroactively clear the models it already trained.
What to watch
- Apple’s fair-use defense as the case moves through discovery on the merits.
- Class certification proceedings following the February 2026 consolidated complaint.
- Whether plaintiffs can obtain internal Apple records detailing exactly which datasets trained Apple Intelligence’s foundation models.
- Coordination or conflicting rulings with other Books3-related litigation against different AI companies.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief