提出基于片段的主动学习方法,用更少标注实现更强的多目标跟踪性能。
Clip-level Uncertainty and Temporal-aware Active Learning for End-to-End Multi-Object Tracking

- 按视频片段而非单帧评估不确定性,匹配端到端追踪器结构
- 仅用50%标注数据即达到全监督效果,提升标注效率
- 适合高成本标注场景下的多目标跟踪系统优化
动态环境中多目标跟踪依赖于可靠的时序推理以维持物体身份一致性。基于Transformer的端到端多目标跟踪模型通过显式建模时序依赖取得优异性能,但其训练需大量边界框和身份标注。鉴于标注成本高且视频中存在强冗余,主动学习(AL)是提升标注效率的有效手段。然而,现有MOT的主动学习方法主要在帧级别操作,与现代端到端追踪器以多帧片段为单位进行推理和训练的结构不匹配。为此,本文提出片段级主动学习框架CUTAL,通过多帧预测的不确定性度量对每个片段评分,捕捉帧间对应关系的模糊性,并引入时序多样性约束,选取信息量丰富且非冗余的样本。实验表明,在MeMOTR和SambaMOTR上,CUTAL在相同标注预算下优于基线方法。值得注意的是,使用仅50%标注数据,CUTAL在两个数据集上对MeMOTR的表现已接近全监督水平。
原文摘要 · Abstract (English)
Multi-Object Tracking (MOT) in dynamic environments relies on robust temporal reasoning to maintain consistent object identities over time. Transformer-based end-to-end MOT models achieve strong performance by explicitly modeling temporal dependencies, yet training them requires extensive bounding-box and identity annotations. Given the high labeling cost and strong redundancy in videos, Active Learning (AL) is an effective approach to improve annotation efficiency. However, existing AL methods for MOT primarily operate at the frame level, which is structurally misaligned with modern end-to-end trackers whose inference and training rely on multi-frame clips. To bridge this gap, we formulate clip-level active learning and propose Clip-level Uncertainty and Temporal-aware Active Learning (CUTAL). In contrast to frame-based approaches, CUTAL scores each clip using uncertainty metrics derived from multi-frame predictions to capture inter-frame correspondence ambiguities, while enforcing temporal diversity to select an informative and non-redundant subset. Experiments show that CUTAL achieves stronger overall performance than baselines at the same label budgets across MeMOTR and SambaMOTR. Notably, CUTAL achieves performance comparable to full supervision for MeMOTR on both datasets using only 50% of the labeled training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。