Glossary
Synthetic data
Data generated by a model rather than recorded from people or the world. It can extend a corpus cheaply, but models trained mostly on model output tend to degrade, so human-made recordings remain the ground truth.
Model-generated examples are cheap to scale and useful for filling gaps, simulating rare cases, and standing in for sensitive records. But training repeatedly on model output can degrade quality across generations — described in research as model collapse — and licences increasingly state whether one model’s outputs may train another. Provenance separates the two: human-made data traces back to a real recording event.
Why it matters
The more synthetic data circulates, the more verified human-origin data functions as ground truth.