Resources
Licensed vs scraped training data: what’s safe to train on
The web is public, so scraping feels free. The law is more complicated. Here is what the current cases actually decided, and what safe to train on means for a buyer.
Published 2026-07-22 · 7 min read
Key takeaways
- Fair use is a defense decided after a lawsuit, not a rule that makes scraping safe.
- The cases people cite as wins are narrower than the headlines. Bartz settled, NYT is ongoing, Ross was non-generative, and Getty largely lost in the UK.
- Scraped data can carry contract, privacy, and biometric exposure that copyright analysis misses entirely.
- Licensed data keeps provenance, consent, and indemnity attached to each asset.
- Non-exclusive, per-asset licences cover most training uses without paying for exclusivity you do not need.
It’s on the internet, so it’s fair game. Most training-data decisions start from that assumption. It does not hold. A public web page is not a public-domain page. Scraping copies protected works, and copyright is only the first of several exposures a buyer inherits.
This article separates settled law from open questions. The short version: no US court has blessed scraping copyrighted work for AI training as a general rule, and the cases people cite as wins are narrower than the headlines suggest. Licensed data is not a legal nicety. It is how a buyer keeps provenance, consent, and indemnity attached to the asset.
The objection: it’s public, so it’s fair game
Fair use is a defense, not a permission slip. It is decided case by case, weighed across the statutory factors, after someone has already sued. Planning a training run around we’ll argue fair use later is planning to litigate.
Public access also does not clear the other rights attached to a recording. A podcast episode on an open page still carries the copyright of its producer, the performances of its speakers, and, where a person is identifiable, voice and likeness interests. Scraping copies all of that silently. Licensing makes each layer explicit.
What the cases have actually decided
Bartz v. Anthropic settled. Anthropic agreed to a payment widely reported as the largest in US copyright history, and the case resolved before any appellate ruling. A settlement binds the parties. It sets no precedent. The dispute also turned heavily on pirated copies, not on a clean holding that training itself is lawful. Read it as a warning about how you acquire data, not a green light.
NYT v. OpenAI is ongoing. It sits in discovery in the Southern District of New York, with fights over training logs and produced samples. Nothing on the merits has been decided. Treat any confident claim about its outcome as speculation.
Thomson Reuters v. Ross is often miscited. Ross built a legal-research search tool, not a generative model. The trial court rejected its fair-use defense for copying Westlaw headnotes, and the case is now on appeal at the Third Circuit, argued in June 2026, with a decision pending. It is a narrow, non-generative dispute. It is not a ruling about training chatbots.
Getty v. Stability in the UK is the one most often described backwards. Getty largely lost. It withdrew its primary training claims because it could not show the training happened in the UK, so the court never decided whether training on copyrighted images infringes. The remaining secondary-infringement claim failed. It is not a rights-holder win on training.
Copyright is not the only exposure
Scraped data can carry liabilities that have nothing to do with copyright. A site’s terms of service may prohibit automated collection, which turns scraping into a contract problem. Personal data pulled from EU sources brings the GDPR into scope. Recordings of identifiable people can implicate biometric law, including Illinois BIPA, where a voiceprint counts as biometric information.
None of these show up in a file name. They travel with the data whether or not the buyer knows. Licensing at the source is where these get named, consented, and documented.
What safe to train on means in practice
Safe is not a vibe. It is a set of artifacts attached to each asset. A signed licence that grants AI-training rights. Separate voice and likeness consent where a person is identifiable. A documented chain of title back to the owner. Provenance a buyer can inspect before committing.
That is the model fiund runs on. Every asset carries a signed AI-training licence. Voice and likeness consent is captured separately where it applies. Owners keep ownership and approve which buyers may license their work. Licences are non-exclusive by default, so a buyer pays for use, not for locking the world out. Provenance is surfaced in diligence rather than promised in a slide.
How buyers reduce exposure
Ask for the chain of title before the data, not after. Confirm the licence scope covers model training, not just internal analysis. Check that consent artifacts exist for identifiable people. Keep the indemnity tied to named provenance, so it means something if a claim arrives.
Non-exclusive licensing is usually enough. Most training uses do not need exclusivity, and paying for it raises cost without reducing risk. The goal is a clean record that survives a discovery request, not a trophy.
Sources
Frequently asked questions
Is scraping public web data legal for AI training?
There is no blanket rule that says yes. Fair use is a defense argued case by case, and scraping can also raise contract, privacy, and biometric issues that copyright analysis ignores. Public access does not clear those rights.
Did Bartz v. Anthropic decide that training is infringement?
No. It settled before any appellate ruling, so it sets no precedent. The dispute turned heavily on pirated copies. Read it as a caution about acquisition, not a ruling on training in general.
Does licensed data remove all legal risk?
No single step removes all risk. Licensing attaches provenance, consent, and indemnity to the asset, which is what a buyer relies on if a claim arrives. It moves you from arguing fair use after the fact to holding a documented right up front.
Related resources
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief