Glossary
Action recognition
Identifying what activity is happening in a video clip, such as cooking, assembling, or lifting. Trained on footage annotated with the actions shown.
Datasets pair clips with verbs from a defined taxonomy; research sets such as Kinetics established the pattern. Fine-grained distinctions — tightening versus loosening — are harder than coarse ones and need denser labels. Inside longer footage, temporal annotation marks where each action starts and ends.
Why it matters
Labelled task footage from real settings — kitchens, workshops, job sites — is what supervised activity models train on.