针对视觉目标跟踪中干扰物导致的误判,提出更智能的记忆模型提升追踪精度。
A Distractor-Aware Memory for Visual Object Tracking with SAM2
- 设计干扰物感知记忆机制,动态过滤干扰帧,增强目标表征
- 在七个基准上超越现有方法,在六个上刷新纪录
- 适用于复杂场景下的高鲁棒性视频目标分割任务
基于记忆的追踪器通过将近期追踪到的帧拼接成记忆缓冲区,并通过注意力机制在当前图像中定位目标。尽管已在多个基准上达到顶尖性能,但近期SAM2的发布使这类方法成为视觉目标追踪领域的焦点。然而,现代追踪器在存在干扰物时仍表现不佳。本文认为需要更复杂的记忆模型,提出一种面向干扰物的新型记忆模型与基于内省的更新策略,同时提升分割准确性和追踪鲁棒性。所提出的追踪器命名为SAM2.1++。此外,还构建了一个新的干扰物蒸馏数据集DiDi,以更深入研究干扰问题。SAM2.1++在七个基准上优于SAM2.1及相关SAM记忆扩展,在六个基准上建立新的最先进水平。
原文摘要 · Abstract (English)
Memory-based trackers are video object segmentation methods that form the target model by concatenating recently tracked frames into a memory buffer and localize the target by attending the current image to the buffered frames. While already achieving top performance on many benchmarks, it was the recent release of SAM2 that placed memory-based trackers into focus of the visual object tracking community. Nevertheless, modern trackers still struggle in the presence of distractors. We argue that a more sophisticated memory model is required, and propose a new distractor-aware memory model for SAM2 and an introspection-based update strategy that jointly addresses the segmentation accuracy as well as tracking robustness. The resulting tracker is denoted as SAM2.1++. We also propose a new distractor-distilled DiDi dataset to study the distractor problem better. SAM2.1++ outperforms SAM2.1 and related SAM memory extensions on seven benchmarks and sets a solid new state-of-the-art on six of them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。