Reference

Glossary.

The vocabulary of AI training data, defined plainly.

Chain of title

The documented trail showing who owned a piece of material and what rights transferred at

Provenance

The verifiable record of where data came from and how it was collected and licensed — the

Rights-cleared

Data for which the owner has granted, in writing, the right to use it for AI training, wit

Voice and likeness consent

Permission from an identifiable person to use their voice or appearance for AI training or

Diarization

The task of determining who spoke when in an audio recording, segmenting speech by speaker

ASR (automatic speech recognition)

Technology that transcribes spoken language into text; trained on audio paired with accura

TTS (text-to-speech)

Technology that synthesizes natural-sounding speech from text; trained on clean, consisten

RLHF

Reinforcement learning from human feedback — training a model using human preference judgm

Egocentric video

First-person video captured from a wearable or head-mounted camera, showing a task from th

World model

A model that learns to predict how an environment evolves, often trained on long, continuo

IMU data

Inertial measurement unit data — motion and orientation streams from accelerometers and gy

Agentic trajectory

A recorded sequence of the steps a human takes to complete a task, used to train AI agents

Phoneme coverage

How completely a speech dataset represents the distinct sound units of a language — import

Crosstalk

Overlapping speech where multiple people talk at once; common in natural conversation and

Code-switching

Alternating between two or more languages within a conversation or sentence; common in mul

Not in the crawl

Material that never entered a public web crawl, so a model has not already seen it — worth

Pretraining

The first, large-scale stage of training, in which a model learns general patterns from br

Fine-tuning

Further training of a pretrained model on a smaller, focused dataset to adapt it to a task

Instruction data

Datasets pairing instructions or prompts with high-quality responses, used in supervised f

Preference data

Model outputs annotated with human judgments of which response is better, gathered as pair

Evaluation set

Data withheld from training and used only to measure model performance. A held-out set is

Benchmark contamination

The leak of test or benchmark material into training data, inflating a model’s measured pe

Data deduplication

Removing exact and near-duplicate items from a dataset before training. Duplicates waste c

Synthetic data

Data generated by a model rather than recorded from people or the world. It can extend a c

Foundation model

A large model pretrained on broad data and adaptable to many downstream tasks through fine

Token

The basic unit a model processes — a fragment of text, or a discrete slice of audio or vid

Data curation

The work of selecting, cleaning, filtering, and documenting data before it is used for tra

Data card (datasheet)

A structured document describing a dataset — its sources, collection methods, licensing st

DPO (direct preference optimization)

A fine-tuning method that trains a model directly on preferred-versus-rejected response pa

RLAIF (reinforcement learning from AI feedback)

An alignment method in which the preference judgments come from a judge model applying wri

Model weights

The numerical parameters a model learns during training — the artifact that constitutes th

Transcription

The written record of what was said in a recording. Accurate transcripts paired with their

Forced alignment

Automatically matching each word or phoneme in a transcript to its exact timestamp in the

Voice cloning

Synthesizing a specific person’s voice from recorded samples of their speech. Lawful cloni

Prosody

The rhythm, stress, and intonation of speech — the qualities that make a voice sound natur

Sample rate

The number of times per second an audio signal is measured, expressed in kilohertz. Podcas

Lossless vs lossy audio

Lossless formats such as WAV and FLAC preserve the full recorded signal; lossy formats suc

WER (word error rate)

The standard measure of transcript accuracy, counting substitutions, insertions, and delet

Acoustic conditions

The sound environment captured in a recording — room tone, reverberation, background noise

Audio stems

The separate unmixed tracks of a recording — for example, one file per speaker. Isolated t

Spontaneous vs read speech

Spontaneous speech is unscripted natural talk, complete with fillers, interruptions, and r

Accent coverage

How well a speech dataset represents the accents and dialects of a language. Broad coverag

Speaker metadata

Structured information about who is speaking — accent, age range, gender, languages, and r

MOS (mean opinion score)

The standard subjective measure of speech and audio quality, in which listeners rate sampl

VAD (voice activity detection)

Identifying which portions of an audio signal contain speech and which are silence, music,

Wake word

A short phrase, such as an assistant’s name, that a device listens for continuously to kno

Far-field audio

Speech recorded at a distance from the microphone — across a room rather than into a heads

SNR (signal-to-noise ratio)

The ratio of desired signal power to background noise power, expressed in decibels. Higher

Motion capture

Recording human or object movement as structured position data using markers, suits, or se

Pose estimation

Detecting the position of a person’s body in images or video as a skeleton of keypoints su

Action recognition

Identifying what activity is happening in a video clip, such as cooking, assembling, or li

Temporal annotation

Labels tied to time ranges in audio or video — when an action starts and ends, who is spea

Dashcam footage

Continuous road video from vehicle-mounted cameras, treated as a distinct data class for t

B-roll

Supplementary footage shot around a main production — establishing shots, cutaways, and pr

Optical flow

The per-pixel motion field between consecutive video frames — the direction and distance e

AI training licence

A written agreement granting the right to use specified material to train, fine-tune, or e

Exclusivity

A licence term controlling whether the same data can be licensed to others. Exclusive gran

Per-licence pricing

Charging per licence granted rather than per unit of data. Because non-exclusive rights ca

Revenue share

An arrangement paying the data owner a percentage of licensing revenue, or a recurring roy

Licence term

The period during which a licence remains in force. Because training embeds data in a mode

Territory

The geographic scope of a licence — where the licensee may use the data or deploy models t

Sublicensing

A licensee’s right to pass licensed data on to third parties such as subsidiaries, contrac

Derivative model rights

Terms governing models trained on licensed data — whether the licensee may keep, commercia

Warranty

A supplier’s contractual assurance that stated facts about the data are true — that they o

Indemnification

A contractual promise to cover the other party’s losses if a specified risk materializes —

Audit rights

A licensor’s right to verify that a licensee is using data only as agreed — for example, c

Takedown and clawback

Contract terms requiring a licensee to stop using and delete data in defined situations, s

Release form

A signed document in which a participant grants defined rights in their recorded voice, im

Work made for hire

A copyright doctrine under which work created by an employee, or under a qualifying writte

PII (personally identifiable information)

Information that can identify a specific person — names, contact details, and under many p

De-identification

Removing or masking details that link data to a specific person, such as bleeping names or

Voiceprint (biometric identifier)

The distinctive, measurable characteristics of a person’s voice, treated as a biometric id

TDM exception (text and data mining)

A statutory copyright exception, notably in the EU, permitting text and data mining of law

AI training opt-out

A machine-readable signal — robots.txt rules, ai.txt files, or embedded metadata — asking

Fair use

A US doctrine permitting limited unlicensed use of copyrighted work, weighed case by case

Public domain

Material in which copyright has expired or never applied, free for anyone to use. Public-d

Creative Commons and AI training

A family of public copyright licences. CC licences predate modern AI training and do not a

DMCA

The US Digital Millennium Copyright Act, known for its notice-and-takedown process and its

Content credentials (C2PA)

An open standard from the C2PA coalition for attaching signed provenance metadata to media

Watermarking

Embedding an imperceptible signal inside media — increasingly inside AI-generated output —

Content fingerprinting

Deriving a compact identifier from a piece of media so copies can be recognized wherever t