Glossary

Evaluation set

Data withheld from training and used only to measure model performance. A held-out set is trustworthy only if the model has never seen it, which is why unexposed material is sought for benchmarks.

Evaluation sets are kept small, stable, and unpublished, because results steer model releases and public claims. Licences for evaluation-only use are narrower than training licences: the data may be run through a model but not trained into it. Never-published material is prized because any exposure to a public crawl undermines the measurement.

Why it matters

The same archive can be licensed for evaluation and for training separately, on different terms.

See also

← All terms