Glossary

Action recognition

Identifying what activity is happening in a video clip, such as cooking, assembling, or lifting. Trained on footage annotated with the actions shown.

Datasets pair clips with verbs from a defined taxonomy; research sets such as Kinetics established the pattern. Fine-grained distinctions — tightening versus loosening — are harder than coarse ones and need denser labels. Inside longer footage, temporal annotation marks where each action starts and ends.

Why it matters

Labelled task footage from real settings — kitchens, workshops, job sites — is what supervised activity models train on.

See also

← All terms