Flagship report

The State of AI Training Data Licensing 2026

What the courts decided, what regulators now require, and how the market for consented, rights-cleared training data took shape this year.

Published 2026-07-22 · Updated 2026-07-22 · 16 min read

Key findings
  • A federal court approved the Bartz v. Anthropic settlement at about $1.5 billion, the largest copyright recovery on record. A class settlement sets no binding precedent, so it does not resolve whether training on copyrighted work is fair use.
  • The first appellate test of AI training fair use, Thomson Reuters v. Ross, was argued at the Third Circuit on June 11, 2026. A decision is pending. Until it lands, the core question stays open.
  • The New York Times v. OpenAI is in discovery in the Southern District of New York, with trial expected in late 2026 or 2027. It remains the most consequential fair-use test in the United States.
  • Getty v. Stability AI is not a rights-holder training win. Getty dropped its UK training claim for lack of jurisdiction, so the court never ruled on whether training itself infringes.
  • The EU AI Act training-data summary duty under Article 53 began applying on August 2, 2025, with a mandatory public template. Provenance disclosure is now a legal obligation for general-purpose model providers.
  • Voice and likeness law is hardening. The Tennessee ELVIS Act is in force, and the federal NO FAKES Act cleared the Senate Judiciary Committee on June 18, 2026.
  • Biometric-privacy exposure is rising. Illinois BIPA voiceprint class actions now target voice-model training directly, including 2026 suits brought by professional voice actors.
  • Disclosed deals set public benchmarks. News Corp and OpenAI was reported above $250M over five years; Reddit and Google at about $60M a year. These read as a market, not a favor.
  • Demand is moving to data the open crawl never held: egocentric, agentic, world-model, far-field, and low-resource multilingual. For those, consented sourcing is the only supply path.

Contents

  1. The year in one page
  2. The legal weather: what courts decided, what stays open
  3. Regulation tightens: disclosure, voice, and biometrics
  4. The licensing market comes of age
  5. The demand shift: past the crawl
  6. What buyers now require in diligence
  7. What owners should do in 2026
  8. Outlook: 2026 into 2027
  9. Sources

The year in one page

This was the year AI training data stopped being a free input and became a priced one. The shift did not come from a single ruling. It came from a stack of them, plus new statutes, plus a run of licensing deals large enough to set public reference points.

The headline number belonged to Bartz v. Anthropic. A federal court approved a settlement of about $1.5 billion, the largest copyright recovery on record. It rewarded authors whose books were pulled from shadow libraries to train a model. It also came with an order to destroy the unlawfully obtained files. What it did not do is decide the law. A class settlement binds the parties, not future courts. The fair-use question it seemed to answer is still open.

That open question now sits with two courts. Thomson Reuters v. Ross was argued at the Third Circuit in June, the first appellate look at whether training on copyrighted work is fair use. The New York Times v. OpenAI is grinding through discovery in New York. Neither has resolved. Buyers who wanted a clean legal answer this year did not get one.

While the courts hesitated, regulators moved. The EU AI Act began requiring general-purpose model providers to publish a summary of their training content, on a template the AI Office set. In the United States, voice and likeness statutes advanced at both the state and federal level, and biometric-privacy suits reached into voice-model training. The direction is consistent. Sourcing has to be documented, and consent has to be real.

The market read the signal. Publishers, forums, and stock libraries signed licensing deals at values large enough to report. And demand began shifting toward data the open web never contained. Taken together, 2026 is the year provenance became a purchase requirement rather than a preference.

Takeaways
  • The largest recovery on record settled without setting precedent.
  • The law is unsettled; disclosure and consent duties are not.
  • Provenance moved from nice-to-have to buy-side requirement.

Regulation tightens: disclosure, voice, and biometrics

Regulation did what litigation would not: it set rules. The clearest is the EU AI Act. Under Article 53, providers of general-purpose AI models must publish a sufficiently detailed summary of the content used to train the model, following a template the AI Office released on July 24, 2025. The obligation began applying on August 2, 2025. Models placed on the market before that date have until August 2, 2027 to comply, and from August 2, 2026 the AI Office can verify and require corrective measures. The template favors aggregated, narrative disclosure over a work-by-work list, and asks providers to describe how they honored text-and-data-mining opt-outs. The point is structural. Provenance is now something a model provider has to write down and stand behind.

In the United States, the action is on voice and likeness. The Tennessee ELVIS Act, the Ensuring Likeness Voice and Image Security Act, took effect on July 1, 2024. It extended the state right of publicity to cover voice explicitly and reached the unauthorized use of tools whose primary purpose is to make an unauthorized digital replica. At the federal level, the NO FAKES Act was advanced by the Senate Judiciary Committee on June 18, 2026. It would create a federal right against unauthorized digital replicas of a person voice or likeness. It has not passed either chamber, and it is drafted to preserve existing state laws such as the ELVIS Act. For anyone licensing real human voices, the consent chain is becoming a statutory requirement, not just a contract nicety.

Biometric privacy is the third front. Illinois BIPA continues to generate voiceprint litigation, and the theory has moved toward AI. In 2026, a group of professional voice actors and narrators brought coordinated class actions alleging that voiceprints were extracted from public recordings and used to train voice models without consent. A separate class was certified in 2025 covering roughly 1.2 million Illinois residents for whom a voiceprint was created. In Europe, GDPR Article 9 treats biometric data processed to uniquely identify a person as a special category, which generally requires explicit consent. Voice is not just content. Where it identifies a person, it is regulated biometric data.

The common thread across all three is documentation. Disclose what trained the model. Prove the voice was licensed. Show the consent that covers the biometric. A dataset that cannot answer those three questions is now a compliance liability, whatever a court eventually says about fair use.

Takeaways
  • EU AI Act Article 53 makes training-data disclosure a legal duty for GPAI providers.
  • The ELVIS Act is live; NO FAKES cleared committee but is not yet law.
  • BIPA and GDPR Article 9 turn unlicensed voice into biometric exposure.

The licensing market comes of age

For years the counterargument to licensing was that there was no market to point to. That excuse is gone. In 2026 the licensing of training data has visible, reported prices and repeat buyers.

The reference points are public. OpenAI multi-year deal with News Corp was reported at more than $250 million over five years, the largest disclosed content-licensing agreement. Reddit licensed its posts to Google for about $60 million a year, and disclosed aggregate data-licensing contracts worth roughly $203 million in its IPO filing. Reddit also struck a comparable arrangement with OpenAI. Stock libraries turned licensing into a line of business: Shutterstock reported growing AI-licensing revenue year over year. None of these are fiund figures. They are market facts, reported in the press and in public filings, and they establish that trained-on data has a clearing price.

Two things about the deal wave matter beyond the numbers. First, the counterparties are exactly the owners of large, rights-managed, human-made corpora: news publishers, community forums, image and video libraries, and platform archives. That is the supply side of a real market forming around content whose provenance can be established. Second, the deals are being renegotiated, not just signed. Reports in mid-2026 that Reddit was weighing whether to renew or restrict Google access show that owners now treat this as recurring, priced access rather than a one-time sale. Leverage is shifting toward whoever can prove clean rights.

Analysts have started to size the category. Estimates vary by definition, but the dataset-licensing-for-AI-training segment has been valued in the low single-digit billions for 2025 with high-teens compound growth projected through the early 2030s, and broader AI-training-dataset market estimates run higher still. The precise figure matters less than the trend. Money that used to be spent on crawling and scraping is being redirected toward licensing, labeling, and provenance.

The signal for owners is simple. If national publishers and top forums can license their archives, so can holders of specialized audio, video, and user-generated material, provided the rights are documented. The signal for buyers is equally simple. Paying for data is now normal, defensible, and increasingly expected in diligence.

Takeaways
  • Reported deals (News Corp, Reddit, Shutterstock) prove a clearing price exists.
  • Owners of rights-managed corpora are the supply side; renewals show recurring value.
  • Spend is moving from scraping toward licensing and provenance.

The demand shift: past the crawl

The most important change in 2026 is not legal. It is where demand is going. The open web has been crawled. The marginal value of one more scrape of text is falling, and the data that improves frontier models increasingly does not exist in any crawl.

The clearest example is egocentric data: the world seen from a first-person point of view, the way a wearable device or a robot would see it. Human first-person video is far cheaper to collect than robot teleoperation and covers a far wider range of tasks and settings, and it transfers to robot policies. 2025 and 2026 saw large releases pushing this into the mainstream of robot learning, and researchers reported scaling laws tying more human egocentric data to better downstream robot performance. This data has to be captured on purpose, from consenting people, in real environments. It cannot be scraped.

The same is true across the frontier modalities. World-model training wants long, continuous, physically grounded video. Agentic training wants real task traces, screens, and workflows rather than static pages. Far-field and multi-speaker audio wants recordings made at distance and in noise, not clean studio takes. Each of these is a category where the only lawful, high-quality supply is sourced and consented.

Language is the other axis. Most speech systems still underperform outside a handful of high-resource languages. The next gains in conversational AI come from dialect, code-switching, and low-resource coverage, where public datasets are thin or absent. Billions of speakers sit in languages that rarely appear in any crawl. Serving them requires collecting new speech from native speakers under proper consent, which is a licensing and sourcing problem before it is a modeling one.

Put the two together and the picture is clear. The data that is abundant is losing value. The data that is scarce, egocentric, agentic, far-field, and low-resource multilingual, is exactly the data that must be sourced from real people with real rights. Scarcity and consent are pushing in the same direction, toward licensed supply.

Takeaways
  • The crawl is exhausted; frontier gains need data that was never online.
  • Egocentric, agentic, world-model, and far-field data must be captured on purpose.
  • Low-resource multilingual speech is a sourcing problem, and sourcing means consent.

What buyers now require in diligence

Buy-side diligence changed this year. A dataset that arrives without a provenance story is no longer cheap. It is a risk that has to be priced, quarantined, or rejected. The questions have become standard.

Where did it come from. Buyers now ask for the origin of every asset, not a vague assurance. Scraped-from-the-open-web is an answer that triggers scrutiny, because it maps directly onto the theories at issue in the pending cases. Sourced, commissioned, or licensed-from-the-owner is the answer that survives review.

Who consented, and to what. For anything with a human in it, a voice, a face, a body, buyers want the consent instrument and its scope. Does it cover AI training. Does it cover derivative and synthetic outputs. Is it revocable, and what happens on revocation. The ELVIS Act, the NO FAKES direction, and BIPA make this the sharpest question for audio and video.

What can the seller actually convey. Owning a recording is not the same as holding the rights to license it for training. Buyers want the chain from the individual, through any platform or aggregator, to the seller, with the training grant intact at every link. Broken chains are where liability hides.

Can it be documented and disclosed. With the EU AI Act summary duty live, buyers increasingly want data they can describe in a public training-content summary without exposure. That favors datasets with clean, aggregable provenance and disfavors mystery corpora. Indemnities are being read more carefully too, since an indemnity from a thinly capitalized seller is not real coverage.

The net effect is a flight to documented supply. Datasets with signed rights, per-licence terms, and a clear consent record command a premium and clear diligence quickly. Everything else waits in legal review or gets cut. Provenance is now the product.

Takeaways
  • Origin, consent scope, chain of rights, and disclosability are now standard asks.
  • Scraped-from-the-web is a risk flag; sourced and licensed clears review.
  • Signed rights and per-licence terms move to the front of the queue.

What owners should do in 2026

If you hold audio, video, or user-generated material, this is a favorable year to be a seller, provided you can prove what you have. The market is paying for exactly what you own, but only once the rights are documented.

Start with an inventory. Know what you hold, in what formats, at what quality, and how much of it there is. Frontier buyers care about volume, diversity, and realism. A large, varied, real-world corpus is worth more than a small clean one, but only if it is organized enough to describe.

Then fix the rights. This is the work that turns a corpus into a licensable asset. For anything with people in it, secure consent that expressly covers AI training and, ideally, synthetic and derivative use. Where you rely on old contracts or platform terms, read them for whether a training grant is actually there. Gaps found now are cheaper than gaps found in a buyer diligence.

Build the provenance record as you go. Capture origin, dates, contributors, consent instruments, and the chain of transfer, and keep it in a form you could hand to a buyer or reflect in a training-content summary. Provenance you can show is the difference between a premium licence and a rejected one.

Decide your licensing posture. Exclusive or non-exclusive. Per-model or blanket. Time-limited with renewal, which the deal wave suggests is where leverage sits. Retain the right to audit use where you can. Owners who treat access as recurring and priced, rather than a one-time sale, are capturing more of the value.

One caution. Do not oversell what you can defend. A signed, narrow, well-documented licence is worth more than a broad grant you cannot stand behind. In a market where buyers are pricing risk, credibility is the asset.

Takeaways
  • Inventory the corpus, then fix the rights before going to market.
  • Secure consent that names AI training and synthetic use.
  • Treat access as recurring, priced, and auditable, not a one-time sale.

Outlook: 2026 into 2027

Expect the legal fog to thin, slowly. The Third Circuit decision in Ross will be the first appellate word on training fair use, and it will shape how buyers and sellers talk about risk even where it is not binding. The New York Times v. OpenAI trial, if it reaches a verdict, will be the first full merits ruling on training and journalism. Neither will settle everything. Both will narrow the range of reasonable positions, which is what a market needs.

Expect disclosure to become routine. The EU AI Act training-content summaries will start appearing at scale, and once a few large providers publish, the format becomes a norm. Buyers will begin asking whether a dataset can be described in one without exposure. That question quietly rewards clean provenance and penalizes mystery corpora.

Expect voice and likeness rules to keep hardening. Whether or not the NO FAKES Act passes in this session, more states will legislate, and biometric-privacy suits will keep testing whether voice-model training crossed a line. The safe posture for anyone touching real human voices is documented, revocable, training-specific consent, and that posture is only going to get more valued.

Expect demand to keep moving toward the scarce and the sourced. Egocentric, agentic, world-model, far-field, and low-resource multilingual data will absorb a growing share of spend, because that is where model gains now come from and because none of it can be scraped. The commodity end of the market, one more copy of the open web, keeps losing value.

The through line for 2027 is that provenance becomes price. The datasets that command premiums will be the ones with signed rights, real consent, and a chain that survives scrutiny. The ones that cannot show their sourcing will keep getting cheaper, or quarantined, or left on the shelf. Licensed, consented, documented supply is the direction of travel, and 2026 is the year it stopped being optional.

Takeaways
  • Ross and the NYT trial will narrow, not end, the legal uncertainty.
  • Training-content disclosure becomes a norm buyers ask for.
  • Provenance becomes price: documented, consented data wins the premium.

Sources

Every figure and case status above is drawn from the public record and press reporting cited here. Deal values are as reported by their sources; settlement and litigation figures are on the public docket.

  1. Kluwer Copyright Blog — The Bartz v. Anthropic Settlement: Understanding America’s Largest Copyright Settlement
  2. Yahoo Finance / Reuters — Anthropic $1.5 billion copyright settlement approved by judge
  3. LawSites — At 3rd Circuit, Judges Press ROSS and Thomson Reuters on Fair Use, AI Training and Market Harm
  4. Baker Botts — Third Circuit Hears Oral Argument in Ross v. Reuters AI Training Copyright Case
  5. LegalClarity — New York Times vs. OpenAI Lawsuit Status and Timeline
  6. Hodder Law — OpenAI, Discovery, and Privacy: What the New York Times Litigation Reveals
  7. Wiggin LLP — Getty Images v Stability AI: High Court delivers landmark judgment
  8. The Register — Getty loses UK copyright battle against Stability AI
  9. Mayer Brown — EU AI Act News: GPAI Rules Start Applying, Training-Data Summary Template Finalized
  10. WilmerHale — European Commission Releases Mandatory Template for Public Disclosure of AI Training Data
  11. EU Artificial Intelligence Act — Overview of Guidelines for GPAI Models
  12. Holland & Knight — Senate Judiciary Committee Advances Legislation to Protect Name, Image, Likeness and Voice (NO FAKES Act)
  13. Wikipedia — ELVIS Act (Ensuring Likeness Voice and Image Security Act)
  14. American Bar Association — Voiceprints, AI, and BIPA: New Trends in Biometric Privacy Litigation
  15. Privacy World — 2025 Year-In-Review: Biometric Privacy Litigation
  16. CBS News — Google strikes $60 million deal with Reddit to train AI models on human posts
  17. Quartz — The price of AI training data, from $5M to $250M
  18. Columbia Journalism Review — Reddit Is Winning the AI Game
  19. CNBC — Reddit stock sinks on report it may not renew Google AI content deal
  20. Grand View Research — AI Training Dataset Market Size & Share Report, 2026-2033
  21. Dataintelo — Dataset Licensing for AI Training Market Research Report
  22. Digital Divide Data — Why Egocentric Datasets Are Becoming The New Standard For Training Robotics Models
  23. Labellerr — 10 Egocentric Datasets Reshaping Robotics and AI in 2026
  24. Surfing Technology — Multilingual ASR and Low-Resource Languages: Why Speech Data Collection Matters More Than Ever

Keep reading

Building or selling training data in 2026?

Whether you are sourcing consented data or licensing an archive you own, the first conversation is short and specific.

Get in touch