arXiv:2506.13457cs.CV2025-06综述被引 10

系统梳理深度学习多目标跟踪方法,从基础到前沿全面解析。

Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art

  • 按检测与关联方式将追踪方法分为五类,结构清晰
  • 2022年新方法显著提升性能,尤其在复杂运动场景
  • 适合研究者快速掌握技术脉络与选型参考

多目标跟踪(MOT)是计算机视觉的核心任务,旨在视频帧中检测物体并跨时间关联。深度学习的兴起极大推动了该领域发展,尤其在基于检测的追踪范式下,该方法仍是主流。2022年,ByteTrack和MOTR等新方法显著加速了进展。本综述深入分析深度学习驱动的MOT方法,系统将基于检测的方法分为五类:联合检测与嵌入、启发式、运动建模、亲和力学习和离线方法。同时,探讨端到端追踪方法,并与传统方案对比。我们在多个基准上评估最新追踪器性能,特别考察其跨领域的泛化能力。结果表明,启发式方法在密集人群且运动呈线性的数据集上表现最优;而基于深度学习的关联方法,无论在基于检测还是端到端框架中,均在复杂运动场景中更优。

原文摘要 · Abstract (English)

Multi-object tracking (MOT) is a core task in computer vision that involves detecting objects in video frames and associating them across time. The rise of deep learning has significantly advanced MOT, particularly within the tracking-by-detection paradigm, which remains the dominant approach. Advancements in modern deep learning-based methods accelerated in 2022 with the introduction of ByteTrack for tracking-by-detection and MOTR for end-to-end tracking. Our survey provides an in-depth analysis of deep learning-based MOT methods, systematically categorizing tracking-by-detection approaches into five groups: joint detection and embedding, heuristic-based, motion-based, affinity learning, and offline methods. In addition, we examine end-to-end tracking methods and compare them with existing alternative approaches. We evaluate the performance of recent trackers across multiple benchmarks and specifically assess their generality by comparing results across different domains. Our findings indicate that heuristic-based methods achieve state-of-the-art results on densely populated datasets with linear object motion, while deep learning-based association methods, in both tracking-by-detection and end-to-end approaches, excel in scenarios with complex motion patterns.

多目标跟踪深度学习综述计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。