提升视觉目标跟踪抗干扰能力,有效防止误追踪相似物体。
Distractor-Aware Memory-Based Visual Object Tracking
- 引入感知干扰物的记忆模块,动态管理记忆信息以减少误跟踪。
- 在13个基准上超越SAM2.1,10个任务创最优性能,遮挡后重检测能力显著提升。
- 适配多种跟踪架构,可实时部署,适合需要高鲁棒性的实际应用。
近期基于记忆的视频分割方法(如SAM2)在分割任务中表现卓越,但在视觉目标跟踪中仍面临干扰物(即与目标外观相似的物体)带来的挑战。本文提出一种干扰物感知的即插即用记忆模块与基于内省的记忆管理方法,构建DAM4SAM。该设计有效抑制了跟踪漂移至干扰物的现象,并提升了目标被遮挡后的重检测能力。为便于分析干扰物影响下的跟踪表现,我们构建了DiDi——一个经过干扰物提炼的数据集。DAM4SAM在13个基准上优于SAM2.1,且在10个任务上达到新基准水平。将该记忆模块集成至实时跟踪器EfficientTAM,性能提升11%,达到非实时版SAM2.1-L的追踪质量;与边缘特征跟踪器EdgeTAM结合,性能提升4%,展现出良好的跨架构泛化能力。
原文摘要 · Abstract (English)
Recent emergence of memory-based video segmentation methods such as SAM2 has led to models with excellent performance in segmentation tasks, achieving leading results on numerous benchmarks. However, these modes are not fully adjusted for visual object tracking, where distractors (i.e., objects visually similar to the target) pose a key challenge. In this paper we propose a distractor-aware drop-in memory module and introspection-based management method for SAM2, leading to DAM4SAM. Our design effectively reduces the tracking drift toward distractors and improves redetection capability after object occlusion. To facilitate the analysis of tracking in the presence of distractors, we construct DiDi, a Distractor-Distilled dataset. DAM4SAM outperforms SAM2.1 on thirteen benchmarks and sets new state-of-the-art results on ten. Furthermore, integrating the proposed distractor-aware memory into a real-time tracker EfficientTAM leads to 11% improvement and matches tracking quality of the non-real-time SAM2.1-L on multiple tracking and segmentation benchmarks, while integration with edge-based tracker EdgeTAM delivers 4% performance boost, demonstrating a very good generalization across architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。