UniTrack通过可微图学习提升多目标跟踪精度,无需改架构即可显著减少身份切换。
UniTrack: Differentiable Graph Representation Learning for Multi-Object Tracking
- 设计可微图损失函数,统一优化检测、身份保持与时空一致性
- 在多个基准上实现最高12% IDF1提升,身份切换减少53%
- 兼容现有追踪模型,尤其适合追求高精度的视频追踪应用
我们提出UniTrack,一种即插即用的图论损失函数,通过统一可微学习直接优化多目标跟踪(MOT)任务的目标,显著提升跟踪性能。与以往需重设计架构的图基方法不同,UniTrack提供通用训练目标,将检测精度、身份保持和时空一致性整合为单一端到端可训练损失,无需修改现有追踪系统架构即可无缝集成。通过可微图表示学习,网络能学习跨帧运动连续性与身份关系的全局表征。我们在多种追踪模型(Trackformer、MOTR、FairMOT、ByteTrack、GTR、MOTE)和多个挑战性基准上验证了UniTrack,所有测试架构均获得一致提升。在复杂基准上,身份切换最多降低53%,IDF1提升12%;其中GTR在SportsMOT上达到9.7% MOTA峰值增益。
原文摘要 · Abstract (English)
We present UniTrack, a plug-and-play graph-theoretic loss function designed to significantly enhance multi-object tracking (MOT) performance by directly optimizing tracking-specific objectives through unified differentiable learning. Unlike prior graph-based MOT methods that redesign tracking architectures, UniTrack provides a universal training objective that integrates detection accuracy, identity preservation, and spatiotemporal consistency into a single end-to-end trainable loss function, enabling seamless integration with existing MOT systems without architectural modifications. Through differentiable graph representation learning, UniTrack enables networks to learn holistic representations of motion continuity and identity relationships across frames. We validate UniTrack across diverse tracking models and multiple challenging benchmarks, demonstrating consistent improvements across all tested architectures and datasets including Trackformer, MOTR, FairMOT, ByteTrack, GTR, and MOTE. Extensive evaluations show up to 53\% reduction in identity switches and 12\% IDF1 improvements across challenging benchmarks, with GTR achieving peak performance gains of 9.7\% MOTA on SportsMOT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。