Multi-object tracking (MOT) provides continuous object localization and identity maintenance in video sequences and supports athlete trajectory extraction, tactical analysis, and automated statistical analysis in sports-video applications1,2,3. Reliable visual information is also related to sport-specific decision-making and performance assessment in sports science, further emphasizing the value of stable video-derived movement data for downstream sports analysis4. In sports scenarios, the reliability of tactical data depends on the continuity of tracking trajectories and the consistency of target identities.
Compared with general MOT scenarios, sports videos contain rapid direction changes, abrupt stops, jumps, dense interactions, and frequent occlusion. These factors cause substantial fluctuations in detection-box confidence and increase the frequency of low-confidence detections, which may interrupt trajectory association5,6. Uniform team jerseys further reduce the discriminative ability of general appearance features and increase identity confusion among visually similar teammates7,8. When occlusion and appearance similarity occur in the same sequence, detection confidence decreases and target appearance information becomes incomplete, increasing the difficulty of trajectory association and identity maintenance9,10. In MOT, re-identification (Re-ID) refers to the process of associating the same target identity across frames, occlusion periods, or re-entry cases using identity-related visual evidence. In team-sport videos, general Re-ID cues may be weakened because athletes from the same team often share similar uniforms and body appearances. The technical literature review indicates that trajectory fragmentation and identity confusion in sports MOT are commonly addressed through two complementary approaches: detection-confidence utilization and identity-cue enhancement. JerseyTrack is positioned at the intersection of these approaches by regulating the use of jersey semantic cues according to detection quality rather than applying color and player-number information uniformly to all detections.
This study presents JerseyTrack, a confidence-guided sports MOT method based on jersey semantic fusion. The method is designed for team-sport videos in which athletes wear visible jerseys and detection results provide bounding boxes with confidence scores. JerseyTrack combines four main components: athlete detection, confidence-level partitioning, jersey semantic feature extraction, and cascaded inter-frame association. Detection confidence is used to adaptively partition detections into high-, medium-, and low-confidence groups, which are processed using different association strategies. High-confidence detections are associated primarily through motion information, whereas medium- and low-confidence detections additionally incorporate jersey color and player-number information to improve identity maintenance when visual evidence is weak. The method is intended for sports video analysis tasks that require stable athlete trajectories, such as tactical analysis and player behavior interpretation.
The proposed approach also has operational limitations. Jersey semantic extraction may be affected by severe occlusion, motion blur, low image resolution, abnormal jersey poses, or invisible player numbers. In these cases, jersey-based semantic cues may become incomplete, and the association process relies more heavily on motion consistency and multi-frame verification. Therefore, the performance claims in this study are restricted to the evaluated datasets and experimental settings. Experiments on the SportsMOT dataset show that JerseyTrack improves the evaluated tracking metrics and reduces identity-switching events under the tested sports tracking conditions. The main difficulty of MOT in sports videos lies in maintaining trajectory continuity and identity consistency under rapid motion, occlusion, and appearance similarity. Existing studies relevant to JerseyTrack can be broadly categorized into two research directions: the use of detection confidence for trajectory association and the enhancement of identity cues through sports-specific semantic features.
Detection confidence has been widely used to estimate the reliability of detection boxes and to guide trajectory association in MOT. Early tracking methods usually treated confidence as a binary filtering criterion, in which detections below a fixed threshold were removed to reduce false positives11,12. This strategy reduces noisy detections but may also discard true targets under occlusion, rapid motion, or low image quality, resulting in trajectory fragmentation13,14,15. Similar filtering strategies have also shown that fixed confidence thresholds are sensitive to scene variation and detection quality16,17. To improve the use of low-confidence detections, ByteTrack introduced an association strategy that retains low-confidence boxes as candidate targets for secondary matching and links them with existing trajectories through motion similarity and trajectory history18. McByte further extends sports MOT association by incorporating temporally propagated segmentation masks as an association cue, improving robustness under fast motion, occlusion, and camera shift conditions without additional per-video training19. Related studies have also adjusted the use of low-confidence detections to retain weak but potentially valid targets during association20,21,22. Confidence-adaptive association strategies subsequently refined matching criteria according to motion consistency and detection reliability23,24,25. Trajectory refinement methods based on detection-box regression and trajectory smoothness optimization have also been used to reduce trajectory fragmentation26,27,28. Additional refinement strategies provide complementary support for maintaining trajectory continuity under complex motion conditions29. Temporal coherent object flow further models object-level temporal consistency across frames, providing another association-oriented approach for reducing trajectory interruption in complex MOT scenarios30. Tracker fusion methods also learn from the outputs of multiple trackers and use learner-based selection or fusion strategies to improve tracking robustness across different MOT benchmarks31. Collectively, these studies indicate that confidence-aware association helps recover interrupted trajectories; however, confidence and motion cues provide limited identity information when players from the same team have similar visual appearances. Although these methods effectively alleviate trajectory fragmentation in general scenarios, they face inherent limitations in sports settings because players on the same team often have highly similar visual appearances, and relying solely on confidence and motion information cannot provide a sufficient basis for identity differentiation32,33,34. Even when low-confidence boxes are effectively recalled, feature similarity can still lead to identity switching, making it difficult to satisfy the stringent identity consistency requirements of sports analysis.
Sports-specific identity cues have also been used to improve identity discrimination in athlete tracking. Jersey semantic information, such as team color and player number, provides identity cues that are more directly related to sports scenarios than general appearance embeddings35,36,37. Jersey color has been used to support team-level discrimination and reduce cross-team confusion, whereas player number recognition provides a finer identity cue when the number region is visible and readable38,39. Existing sports tracking studies have incorporated domain-specific visual features and jersey color recognition to improve player association under crowded sports conditions40,41. Related work on jersey number recognition and synthetic number data further supports identity modeling under low-resolution or partially occluded conditions42,43. These studies indicate that jersey semantics can strengthen identity representation in sports MOT. However, most semantic-enhanced tracking methods do not sufficiently model the relationship between detection confidence and semantic feature quality. High-confidence detections may already contain reliable spatial and appearance information, making repeated semantic extraction less efficient. Low-confidence detections often suffer from occlusion, blur, truncation, or low resolution, reducing the reliability of color and player-number cues. This limitation motivates a confidence-guided association design in which jersey semantics are introduced according to detection quality rather than applied uniformly to all detections. The reviewed studies show that confidence-aware association and jersey semantic modeling address different aspects of sports MOT. Confidence-based strategies mainly determine whether a detection should participate in association and how it should be matched with existing trajectories, whereas jersey semantic strategies primarily enhance identity representation through sports-specific cues. However, these two mechanisms are often designed as separate components. As a result, semantic cues are not always adjusted according to detection quality, and confidence levels do not directly regulate the use of color and player-number information during association. This separation limits the joint handling of trajectory fragmentation and identity confusion in team-sport videos. The remaining research gap lies in linking detection-confidence quality assessment with jersey semantic association within the same tracking workflow.