Resources
The data not in the crawl
Models have already read the open web. The signal that moves them now is the material that never entered a crawl. Here is why not-in-the-crawl data is worth sourcing.
Published 2026-07-22 · 5 min read
Key takeaways
- Frontier models have already trained on the reachable web, so more of it adds little new signal.
- A rising share of the open web is machine-generated, which makes re-scraping it a way to train on model output.
- Not-in-the-crawl material was never exposed to a public crawler, so it sits outside what the model has already learned.
- This data carries both novelty and a known origin, which the open web cannot provide.
- Recording to a brief fills specific gaps deliberately, with consent and provenance attached.
Frontier models have already read the public web. The common crawls, the scraped archives, the open repositories. If a page could be reached and copied, it has likely been in a training set already. That is the problem hiding inside just use more data. More of the same web adds little.
The material that still moves a model is the material that was never in a crawl. Private archives. Unpublished recordings. Consented conversation captured to a brief. It is harder to get, which is exactly why it carries signal the open web cannot.
Why more web data stops helping
A crawl captures what is public and reachable. Once a model has trained on that, feeding it a second copy teaches nothing new. The distribution is already learned. Worse, a growing share of the open web is now machine-generated, so scraping it again risks training on model output rather than human signal.
The scarce resource is not volume. It is novelty and ground truth. Real human speech, real edge cases, real messiness that no one bothered to publish.
What not in the crawl actually is
Not-in-the-crawl is not a synonym for rare. It is material that was never exposed to a public crawler in the first place. A studio’s raw session tapes. A radio archive that predates the web. Field recordings on a hard drive. Multi-speaker conversations recorded with consent for this purpose.
This is where accents, dialects, overlapping speech, background noise, and long-form structure live. The signal a model needs to handle the real world, rather than the tidy version the open web already taught it.
Why it carries more signal
A model learns the tails of a distribution from examples it has not already seen. Not-in-the-crawl material is, by definition, outside the set the model has memorized. It fills gaps the open web left: quiet languages, specific domains, natural conversation that no publisher cleaned up.
It also comes with something scraped data cannot: a known origin. When a recording is sourced to a brief, you know who spoke, in what setting, and under what consent. That provenance is part of the value, not an add-on.
The fiund thesis
fiund exists to move this material into training sets with rights attached. Owners license real audio, video, and user-generated content they already hold or can record. Each asset carries a signed AI-training licence. Where a speaker is identifiable, voice and likeness consent is captured separately. Owners keep ownership and approve buyers.
The result is data a public crawl could never produce: new signal, documented provenance, and consent at the source. Buyers get material that is both novel and defensible. Owners get paid for work that was sitting unused.
Sourced to a brief
The strongest not-in-the-crawl data is often recorded on purpose. A buyer describes what a model is missing: a language, an acoustic setting, a conversational format. Suppliers record to that brief. The gap gets filled deliberately rather than scavenged.
This is the opposite of scraping. Nothing is taken without permission. The material is created or contributed knowingly, priced per licence, and delivered with its origin intact.
Sources
Frequently asked questions
Does not-in-the-crawl just mean rare or obscure data?
No. It means material a public web crawler never reached, whatever its subject. Private archives, raw session recordings, and consented conversation all qualify, even when the topic is ordinary.
Why does novelty matter more than volume?
A model already learned the public web. A second copy of it teaches nothing. Gains come from examples outside the set the model has memorized, which is where not-in-the-crawl material sits.
Related resources
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief