We should implement a form of action and gesture recognition to the toolbox. From our terminology, motion is the continuous displacement in space over time. Actions are motion segments, often (but not always) with a clear beginning and end. Actions can be goal-directed but not always (eg in dance). In music, Actions are often sound-producing.
Gestures are actions with some kind of meaning. The meaning can be direct and connected to natural language, but the meaning can also be expressive, such as in music and dance.
We should implement a multilayered approach in MGT, where actions are derived from motion, and gesture is added to actions (not all).
It is important that action and gesture recognition is based on human activity in a video, not from camera motion or scene change.
We should implement a form of action and gesture recognition to the toolbox. From our terminology, motion is the continuous displacement in space over time. Actions are motion segments, often (but not always) with a clear beginning and end. Actions can be goal-directed but not always (eg in dance). In music, Actions are often sound-producing.
Gestures are actions with some kind of meaning. The meaning can be direct and connected to natural language, but the meaning can also be expressive, such as in music and dance.
We should implement a multilayered approach in MGT, where actions are derived from motion, and gesture is added to actions (not all).
It is important that action and gesture recognition is based on human activity in a video, not from camera motion or scene change.