通过扰动轨迹状态检测关联不稳定性,提升视频片段选择效率。
Probing Association Instability with Track-State Perturbations for Clip-Level Active Learning in Query-Propagation Multi-Object Tracking

- 用双侧扰动追踪状态,衡量预测差异以识别关联不稳定的片段。
- 在DanceTrack和SportsMOT上,相同标注预算下性能优于基线方法。
- 适合需要减少人工标注成本的多目标跟踪主动学习场景。
训练端到端查询传播多目标跟踪(MOT)模型需密集的边界框与身份标注,导致数据构建成本高昂。片段级主动学习通过选择待标注视频片段降低开销,但以往基于输出层时间不确定性的选取标准可能忽略由传播轨迹状态关联不稳定性带来的信息量。本文提出QPID(Query-Propagation Instability and Diversity),一种面向查询传播MOT的片段选取方法,聚焦于传播轨迹状态中的关联不稳定性。QPID通过双向扰动内部轨迹状态并测量与干净参考分支的预测差异来估计该不稳定性。核心思想是:稳定片段中,每个传播轨迹在小扰动下仍应持续追踪同一目标;而模糊片段中,轨迹状态微小变化会导致其追踪目标改变,从而引发定位或置信度变化。QPID采用两种度量:定位漂移与熵加权置信度差异,聚合为片段级关联不稳定性得分。为避免仅因不确定性重复选择,QPID利用不确定性加权视觉覆盖与轨迹级视觉原型,从高不稳定性片段中选出代表性标注批次。在DanceTrack和SportsMOT数据集上,使用MeMOTR和SambaMOTR模型的实验表明,QPID在相同标注预算下表现显著优于现有主动学习基线。
原文摘要 · Abstract (English)
Training query-propagation end-to-end multi-object tracking (MOT) models requires dense bounding-box and identity annotations across video sequences, making dataset construction expensive. Clip-level active learning reduces this cost by selecting video clips for annotation, but prior acquisition criteria based on output-level temporal uncertainty may miss clips whose informativeness comes from association instability in propagated track states. We propose QPID (Query-Propagation Instability and Diversity), a clip acquisition method for query-propagation MOT that targets association instability in propagated track states. QPID estimates this instability by applying two-sided perturbations to internal track states and measuring prediction differences from a clean reference branch. The key idea is that, in stable clips, each propagated track should continue to follow the same target under small perturbations, whereas in ambiguous clips, small changes in the track state can alter which target the track follows, leading to changes in localization or confidence. QPID measures these perturbation-induced prediction differences with two metrics: Localization Drift and Entropy-Weighted Confidence Discrepancy. These metrics are aggregated into a clip-level association-instability score. To avoid redundant uncertainty-only selection, QPID selects a representative annotation batch from high-instability clips using Uncertainty-Weighted Visual Coverage with track-level visual prototypes. Experiments on DanceTrack and SportsMOT with MeMOTR and SambaMOTR show that QPID achieves strong performance compared with active learning baselines under the same annotation budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。