arXiv:2505.18111cs.CV2025-05中稿 · ICPR Multi-Modal V…

将SAM2改进用于多模态目标跟踪,获ICPR挑战赛第一

Adapting SAM 2 for Visual Object Tracking: 1st Place Solution for MMVPR Challenge Multi-Modal Tracking

  • 基于SAM2构建跟踪框架,融合多模态输入增强鲁棒性
  • 在ICPR 2024多模态跟踪挑战中取得89.4的AUC领先成绩
  • 适合关注多模态视觉跟踪与大模型迁移应用的研究者

本文提出一种有效方法,将分割一切模型2(SAM2)适配至视觉目标跟踪(VOT)任务。该方法充分利用SAM2强大的预训练能力,并引入若干关键技术以提升其在VOT中的表现。通过结合SAM2与所提出的优化策略,在2024 ICPR多模态目标跟踪挑战中取得了89.4的AUC得分,位居榜首,验证了该方法的有效性。本文详述了具体方法、对SAM2的关键改进,以及在多模态数据集背景下对结果的全面分析。

原文摘要 · Abstract (English)

We present an effective approach for adapting the Segment Anything Model 2 (SAM2) to the Visual Object Tracking (VOT) task. Our method leverages the powerful pre-trained capabilities of SAM2 and incorporates several key techniques to enhance its performance in VOT applications. By combining SAM2 with our proposed optimizations, we achieved a first place AUC score of 89.4 on the 2024 ICPR Multi-modal Object Tracking challenge, demonstrating the effectiveness of our approach. This paper details our methodology, the specific enhancements made to SAM2, and a comprehensive analysis of our results in the context of VOT solutions along with the multi-modality aspect of the dataset.

目标跟踪SAM2多模态视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。