首次量化分析物体相似性干扰,提出用大模型提升跟踪鲁棒性。
SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
- 用可控实验量化相似物体干扰对追踪的影响
- 引入语义引导后,追踪性能提升最高达0.93 AUC
- 构建首个针对相似干扰的语义引导基准数据集
本文首次系统研究并量化了单目标追踪中的相似物体干扰(SOI)问题。通过受控的在线干扰掩码实验,证实消除干扰源可使所有主流追踪器性能显著提升,最大AUC提升达4.35,直接证明SOI是影响追踪鲁棒性的核心瓶颈,并验证外部认知引导的可行性。基于此,我们提出SOIBench——首个专为应对SOI设计的语义认知引导基准数据集,通过多追踪器集体判断自动挖掘干扰帧,并采用多层级标注协议生成精准语义引导文本。在该数据集上评估发现,现有视觉语言追踪方法无法有效利用语义引导,性能仅小幅提升或反而下降(AUC变化-0.26至+0.71)。为此,我们提出新范式:将大规模视觉语言模型作为外部认知引擎,无缝集成到任意RGB追踪器中,在语义引导下实现显著性能提升(最高AUC增益0.93),大幅超越现有方法。期望SOIBench能成为推动语义认知追踪研究的标准评估平台。
原文摘要 · Abstract (English)
In this paper, we present the first systematic investigation and quantification of Similar Object Interference (SOI), a long-overlooked yet critical bottleneck in Single Object Tracking (SOT). Through controlled Online Interference Masking (OIM) experiments, we quantitatively demonstrate that eliminating interference sources leads to substantial performance improvements (AUC gains up to 4.35) across all SOTA trackers, directly validating SOI as a primary constraint for robust tracking and highlighting the feasibility of external cognitive guidance. Building upon these insights, we adopt natural language as a practical form of external guidance, and construct SOIBench-the first semantic cognitive guidance benchmark specifically targeting SOI challenges. It automatically mines SOI frames through multi-tracker collective judgment and introduces a multi-level annotation protocol to generate precise semantic guidance texts. Systematic evaluation on SOIBench reveals a striking finding: existing vision-language tracking (VLT) methods fail to effectively exploit semantic cognitive guidance, achieving only marginal improvements or even performance degradation (AUC changes of -0.26 to +0.71). In contrast, we propose a novel paradigm employing large-scale vision-language models (VLM) as external cognitive engines that can be seamlessly integrated into arbitrary RGB trackers. This approach demonstrates substantial improvements under semantic cognitive guidance (AUC gains up to 0.93), representing a significant advancement over existing VLT methods. We hope SOIBench will serve as a standardized evaluation platform to advance semantic cognitive tracking research and contribute new insights to the tracking research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。