Glossary
Pose estimation
Detecting the position of a person’s body in images or video as a skeleton of keypoints such as joints. Models learn it from footage annotated with body positions.
Models output keypoints per person per frame — the widely used COCO convention marks 17 body joints — in 2D or 3D. Training and evaluation need footage with annotated joints, and accuracy is scored by comparing predicted to labelled positions. Occlusion, loose clothing, and unusual viewpoints remain the persistent failure cases.
Why it matters
Varied footage of real bodies doing real tasks feeds every downstream motion application.