Resources

Synthetic vs real training data: when real wins

Synthetic data is cheap and endless. It is also a copy of what a model already knows. Here is where synthetic works, and where licensed real data is the only fix.

Published 2026-07-22 · 6 min read

Key takeaways

  1. Synthetic data is a strong tool for balancing, augmentation, format variation, and privacy-safe structure.
  2. It cannot add signal a model has never seen, because it is generated from what the model already encodes.
  3. Training recursively on generated data causes model collapse, where rare cases disappear and quality degrades, as shown in the Nature model-collapse research.
  4. Real, licensed data grounds a model in dialects, natural conversation, and domain audio a generator cannot invent.
  5. The right approach is a mix: synthetic for structure, licensed real data for grounding and the long tail.

Synthetic data is attractive for obvious reasons. It is cheap, fast, and free of the rights questions that come with real recordings. For some jobs it is genuinely the right tool. The mistake is treating it as a replacement for real-world signal.

Synthetic data is generated by a model. That means it can only recombine what a model already encodes. It cannot introduce ground truth the model has never seen. Lean on it too hard and the model drifts. This article draws the line between the two.

What synthetic data is good at

Synthetic data shines when you need structure, not new signal. Balancing a class that is under-represented. Generating format variations. Producing rare combinations of things the model already understands. Stress-testing a pipeline. Protecting privacy by avoiding real personal data where a real person is not required.

In these cases the goal is coverage and control, not novelty. Synthetic data supplies both cheaply. Used this way, it is a legitimate part of a training mix.

The failure mode: model collapse

The risk shows up when models train on their own output, generation after generation. Researchers call the result model collapse. In work published in Nature in 2024, Shumailov and colleagues showed that training recursively on generated data causes the tails of the distribution to disappear. Rare events vanish first, then quality degrades across the board.

This matters now because the open web is filling with machine-generated text and media. Scraping it again is a quiet form of the same recursion. The fix is not more synthetic data. It is fresh, real, human signal that anchors the model to the world.

Grounding: why real data anchors a model

Real recordings carry things a generator cannot invent from within its own distribution. The specific way a dialect clips a vowel. The overlap when two people talk at once. Room noise, breath, hesitation, the texture of an actual voice. This is grounding: contact with the real world that keeps a model honest.

Synthetic speech modeled on a model tends toward the average. Real speech carries the tails. For anything that has to work in the messy real world, the tails are the point.

When real is the only fix

Reach for licensed real data when the target is authenticity and coverage the model does not already hold. New languages and dialects. Natural conversation with real turn-taking. Domain audio a generator has never heard. Anything where the model must match reality rather than its own prior.

These gaps cannot be filled by generating more from the same model. That only re-serves what the model already knows. The only source of genuinely new signal is real material the model has not seen, captured with rights attached.

Using both well

This is not synthetic versus real as a loyalty test. It is a mix. Use synthetic data for balance, augmentation, and privacy-safe structure. Use licensed real data for grounding, novelty, and the long tail. Keep provenance on the real portion so the mix stays defensible.

The teams that get this wrong treat synthetic as free scale. The teams that get it right treat it as a multiplier on a real foundation.

Watch outA rising share of the open web is now machine-generated. Re-scraping it is a subtle form of training on model output, with the same collapse risk. Fresh real data is the anchor against it.

Sources

← All resources

Frequently asked questions

Can synthetic data replace real training data?

Not for signal a model does not already have. Synthetic data is generated from the model, so it recombines what the model knows. It is good for balance and augmentation, but it cannot introduce new ground truth. That requires real data.

What is model collapse?

It is the degradation that happens when models train on generated data across successive generations. The rare tails of the distribution disappear first, then overall quality falls. The Nature research by Shumailov and colleagues documented the effect.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief