Resources
Red flags when buying training data
The warning signs that a dataset will cost you later: too clean, no consent records, vague sourcing, and more.
Published 2026-07-22 · 7 min read
Key takeaways
- Impressive is not the same as safe. The problems surface after training.
- Missing consent for identifiable people is a walk-away, not a discount.
- Adjectives are not sourcing. Ask for the mechanism.
- Audit a random slice, never just the curated sample.
- Unsettled law is why documented, licensed data lowers your risk.
Most bad training datasets do not look bad. They look impressive. The sample is crisp, the deck is clean, and the problems stay hidden until you have already trained on them.
These are the red flags experienced buyers treat as a reason to slow down, ask harder questions, or walk. None of them is proof of a problem on its own. Together, they are a pattern.
Too clean to be true
Real-world audio and video is messy. It has room noise, uneven levels, accents, interruptions, and gaps. A dataset that arrives perfectly balanced, perfectly labelled, and free of any artifact deserves a question. Where did the messiness go.
Sometimes the answer is fine. It was cleaned, and the vendor can show the raw source. Sometimes it is not. The set is partly synthetic, or it was scraped and scrubbed, or the sample was hand-picked and does not represent the delivery. Ask which.
No consent records for identifiable people
This is the hard stop. If people are identifiable in the data, you need their consent for voice or likeness to be used in training, separate from the owner’s licence. A vendor who cannot produce consent records is selling you risk, not data.
Watch for the deflection that the people signed a release. Which release. For what use. AI training is a specific use, and older releases rarely name it. If the consent does not cover training, it does not cover you.
Vague sourcing
Ask where the data came from. A clean vendor gives a specific answer. Commissioned in these conditions. Contributed by these creators. Drawn from this archive. A vendor who answers only with adjectives, "premium", "proprietary", "ethically sourced", and no mechanism is hiding the mechanism.
Vague sourcing usually means a vague chain of title. If they will not tell you how they got it, assume they cannot prove they can sell it.
The one-sample dazzle
One stunning sample proves that one sample is good. It proves nothing about the other thousands of files you are paying for. Be wary of a pitch built entirely around a single clip.
Insist on auditing a random slice that you select, not the reel the vendor curated. Quality and rights both have to hold across the set, not at the top of it.
No chain of title
Chain of title is the unbroken line from the creator to the party selling to you. Break any link and the grant may not be theirs to give. Ask them to walk the line. If they cannot, the licence sits on sand.
Over-broad reuse claims
Be suspicious of a vendor who claims they can grant everything. Rights have limits. A grant that is silent on those limits, or that claims worldwide, all-media, perpetual rights over content full of identifiable people with no consent layer, is either careless or misrepresented.
The law here is unsettled, which is the point. Early rulings have not handed AI training a clear pass. In the United States, one court rejected a fair use defence in a non-generative case and stressed that limit. In the United Kingdom, the Getty v. Stability case ended largely against the rights holder on the core copyright and training questions, and left the big questions open. Licensed, documented data is how you stay out of that fight.
Red flags at a glance
- A dataset too clean to match real-world capture, with no raw source shown
- No consent records for identifiable people, or releases that never name training
- Sourcing described in adjectives, not mechanisms
- A pitch built on one sample, with no random audit allowed
- A chain of title the vendor cannot walk end to end
- Reuse claims broader than any owner could plausibly grant
Sources
Frequently asked questions
Is scraped data actually illegal to train on?
It is unsettled, not clearly safe. Courts have started to rule, and they have not given AI training a blanket pass. That uncertainty is a business risk you carry into every model trained on undocumented data. Licensed data with provenance removes the question.
A vendor says their data is ethically sourced. Is that enough?
No. Ethically sourced is a claim, not a record. Ask what it means in practice: who consented, how, and to what. Ask to see the documents. A real practice can be shown. A slogan cannot.
What if only some assets have consent gaps?
Then only some assets are usable, and you need to know which. Ask the vendor to segment the set by consent and rights status. A vendor who cannot segment it does not have the records to back any of it.
Related resources
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief