Training begins with researcher-provided labels that mark selected body parts in a small set of video frames. A neural network uses those examples to learn visual features associated with each body part, then predicts their coordinates in new frames. This learned relationship allows the system to follow movement across changing poses and scenes without requiring a physical marker on the animal.
Variation in backgrounds and poses tests whether the learned visual features generalize beyond the frames used for labeling. When the training examples represent different appearances and configurations, predictions can remain useful across more of the recorded behavior. This flexibility is important in biology because animals may move naturally rather than maintaining standardized positions during observation.
Markerless tracking obtains body-part coordinates from visual information in video, whereas physical-marker approaches depend on attaching or placing markers on the subject. Avoiding markers supports observation of movement without that added intervention and can facilitate analysis of natural behavior. The resulting coordinates provide a basis for measuring locomotion, biomechanics, and other time-resolved biological movements.
A typical workflow starts by selecting body parts relevant to the biological question, labeling those parts in a small set of video frames, and using the labels to train a neural network. The trained model then predicts body-part coordinates across new frames. Researchers can use these time-resolved coordinates to quantify movement rather than relying only on visual inspection.
The predicted coordinates convert video into time-resolved movement data for selected body parts. Researchers can use these data to examine animal behavior, locomotion, and biomechanics, including how body parts change position during movement. Because the measurements are generated across video frames, they support quantitative comparisons and more reproducible analysis than qualitative descriptions alone.
The toolkit is useful when researchers need to quantify animal movement across many video frames or study behavior without attaching physical markers. Applications described for biology include automated analysis of behavior, locomotion, biomechanics, and neural function. By turning recorded movement into structured measurements, it can support large-scale studies of natural movement and improve behavioral reproducibility.