Bounding boxes, instance segmentation masks, object and action labels, 21-point hand pose, hand-object interaction labels, task-state labels, failure-event markers, temporal sequence segmentation, and environment metadata — applied per frame, per clip, or per task as your spec requires.