Reference
Glossary.
The vocabulary of AI training data, defined plainly.
Chain of title
The documented trail showing who owned a piece of material and what rights transferred at …
Provenance
The verifiable record of where data came from and how it was collected and licensed — the …
Rights-cleared
Data for which the owner has granted, in writing, the right to use it for AI training, wit…
Voice and likeness consent
Permission from an identifiable person to use their voice or appearance for AI training or…
Diarization
The task of determining who spoke when in an audio recording, segmenting speech by speaker…
ASR (automatic speech recognition)
Technology that transcribes spoken language into text; trained on audio paired with accura…
TTS (text-to-speech)
Technology that synthesizes natural-sounding speech from text; trained on clean, consisten…
RLHF
Reinforcement learning from human feedback — training a model using human preference judgm…
Egocentric video
First-person video captured from a wearable or head-mounted camera, showing a task from th…
World model
A model that learns to predict how an environment evolves, often trained on long, continuo…
IMU data
Inertial measurement unit data — motion and orientation streams from accelerometers and gy…
Agentic trajectory
A recorded sequence of the steps a human takes to complete a task, used to train AI agents…
Phoneme coverage
How completely a speech dataset represents the distinct sound units of a language — import…
Crosstalk
Overlapping speech where multiple people talk at once; common in natural conversation and …
Code-switching
Alternating between two or more languages within a conversation or sentence; common in mul…
Not in the crawl
Material that never entered a public web crawl, so a model has not already seen it — worth…
Pretraining
The first, large-scale stage of training, in which a model learns general patterns from br…
Fine-tuning
Further training of a pretrained model on a smaller, focused dataset to adapt it to a task…
Instruction data
Datasets pairing instructions or prompts with high-quality responses, used in supervised f…
Preference data
Model outputs annotated with human judgments of which response is better, gathered as pair…
Evaluation set
Data withheld from training and used only to measure model performance. A held-out set is …
Benchmark contamination
The leak of test or benchmark material into training data, inflating a model’s measured pe…
Data deduplication
Removing exact and near-duplicate items from a dataset before training. Duplicates waste c…
Synthetic data
Data generated by a model rather than recorded from people or the world. It can extend a c…
Foundation model
A large model pretrained on broad data and adaptable to many downstream tasks through fine…
Token
The basic unit a model processes — a fragment of text, or a discrete slice of audio or vid…
Data curation
The work of selecting, cleaning, filtering, and documenting data before it is used for tra…
Data card (datasheet)
A structured document describing a dataset — its sources, collection methods, licensing st…
DPO (direct preference optimization)
A fine-tuning method that trains a model directly on preferred-versus-rejected response pa…
RLAIF (reinforcement learning from AI feedback)
An alignment method in which the preference judgments come from a judge model applying wri…
Model weights
The numerical parameters a model learns during training — the artifact that constitutes th…
Transcription
The written record of what was said in a recording. Accurate transcripts paired with their…
Forced alignment
Automatically matching each word or phoneme in a transcript to its exact timestamp in the …
Voice cloning
Synthesizing a specific person’s voice from recorded samples of their speech. Lawful cloni…
Prosody
The rhythm, stress, and intonation of speech — the qualities that make a voice sound natur…
Sample rate
The number of times per second an audio signal is measured, expressed in kilohertz. Podcas…
Lossless vs lossy audio
Lossless formats such as WAV and FLAC preserve the full recorded signal; lossy formats suc…
WER (word error rate)
The standard measure of transcript accuracy, counting substitutions, insertions, and delet…
Acoustic conditions
The sound environment captured in a recording — room tone, reverberation, background noise…
Audio stems
The separate unmixed tracks of a recording — for example, one file per speaker. Isolated t…
Spontaneous vs read speech
Spontaneous speech is unscripted natural talk, complete with fillers, interruptions, and r…
Accent coverage
How well a speech dataset represents the accents and dialects of a language. Broad coverag…
Speaker metadata
Structured information about who is speaking — accent, age range, gender, languages, and r…
MOS (mean opinion score)
The standard subjective measure of speech and audio quality, in which listeners rate sampl…
VAD (voice activity detection)
Identifying which portions of an audio signal contain speech and which are silence, music,…
Wake word
A short phrase, such as an assistant’s name, that a device listens for continuously to kno…
Far-field audio
Speech recorded at a distance from the microphone — across a room rather than into a heads…
SNR (signal-to-noise ratio)
The ratio of desired signal power to background noise power, expressed in decibels. Higher…
Motion capture
Recording human or object movement as structured position data using markers, suits, or se…
Pose estimation
Detecting the position of a person’s body in images or video as a skeleton of keypoints su…
Action recognition
Identifying what activity is happening in a video clip, such as cooking, assembling, or li…
Temporal annotation
Labels tied to time ranges in audio or video — when an action starts and ends, who is spea…
Dashcam footage
Continuous road video from vehicle-mounted cameras, treated as a distinct data class for t…
B-roll
Supplementary footage shot around a main production — establishing shots, cutaways, and pr…
Optical flow
The per-pixel motion field between consecutive video frames — the direction and distance e…
AI training licence
A written agreement granting the right to use specified material to train, fine-tune, or e…
Exclusivity
A licence term controlling whether the same data can be licensed to others. Exclusive gran…
Per-licence pricing
Charging per licence granted rather than per unit of data. Because non-exclusive rights ca…
Revenue share
An arrangement paying the data owner a percentage of licensing revenue, or a recurring roy…
Licence term
The period during which a licence remains in force. Because training embeds data in a mode…
Territory
The geographic scope of a licence — where the licensee may use the data or deploy models t…
Sublicensing
A licensee’s right to pass licensed data on to third parties such as subsidiaries, contrac…
Derivative model rights
Terms governing models trained on licensed data — whether the licensee may keep, commercia…
Warranty
A supplier’s contractual assurance that stated facts about the data are true — that they o…
Indemnification
A contractual promise to cover the other party’s losses if a specified risk materializes —…
Audit rights
A licensor’s right to verify that a licensee is using data only as agreed — for example, c…
Takedown and clawback
Contract terms requiring a licensee to stop using and delete data in defined situations, s…
Release form
A signed document in which a participant grants defined rights in their recorded voice, im…
Work made for hire
A copyright doctrine under which work created by an employee, or under a qualifying writte…
PII (personally identifiable information)
Information that can identify a specific person — names, contact details, and under many p…
De-identification
Removing or masking details that link data to a specific person, such as bleeping names or…
Voiceprint (biometric identifier)
The distinctive, measurable characteristics of a person’s voice, treated as a biometric id…
TDM exception (text and data mining)
A statutory copyright exception, notably in the EU, permitting text and data mining of law…
AI training opt-out
A machine-readable signal — robots.txt rules, ai.txt files, or embedded metadata — asking …
Fair use
A US doctrine permitting limited unlicensed use of copyrighted work, weighed case by case …
Public domain
Material in which copyright has expired or never applied, free for anyone to use. Public-d…
Creative Commons and AI training
A family of public copyright licences. CC licences predate modern AI training and do not a…
DMCA
The US Digital Millennium Copyright Act, known for its notice-and-takedown process and its…
Content credentials (C2PA)
An open standard from the C2PA coalition for attaching signed provenance metadata to media…
Watermarking
Embedding an imperceptible signal inside media — increasingly inside AI-generated output —…
Content fingerprinting
Deriving a compact identifier from a piece of media so copies can be recognized wherever t…