arXiv:2409.16902cs.CVcs.AI2024-09中稿 · CVPR被引 9

首个水下伪装目标跟踪数据集,推动视觉语言模型在复杂水下场景的应用。

Underwater Camouflaged Object Tracking Meets Vision-Language SAM2

  • 构建多模态水下伪装目标跟踪数据集UW-COT220
  • VL-SAM2在水下与开放场景中均达顶尖性能
  • 首次系统评估SAM2在珊瑚礁等复杂水下环境的表现

过去十年间,视觉目标跟踪取得显著进展,主要得益于大规模数据集的出现。然而,这些数据集主要集中于开放空气场景,对水下动物追踪——尤其是伪装海洋生物带来的复杂挑战——关注不足。为此,本文提出首个大规模多模态水下伪装目标跟踪数据集UW-COT220。基于该数据集,首次全面评估了当前先进的视觉跟踪方法,包括基于SAM和SAM2的追踪器,在珊瑚礁等复杂水下环境中的表现。结果表明,SAM2相比SAM在处理水下伪装目标方面具有更强能力。此外,本文提出一种基于视频基础模型SAM2的新型视觉-语言跟踪框架VL-SAM2。大量实验验证了其在水下及开放空气场景数据集上均达到先进水平。数据集与代码已公开于https://github.com/983632847/Awesome-Multimodal-Object-Tracking。

原文摘要 · Abstract (English)

Over the past decade, significant progress has been made in visual object tracking, largely due to the availability of large-scale datasets. However, these datasets have primarily focused on open-air scenarios and have largely overlooked underwater animal tracking-especially the complex challenges posed by camouflaged marine animals. To bridge this gap, we take a step forward by proposing the first large-scale multi-modal underwater camouflaged object tracking dataset, namely UW-COT220. Based on the proposed dataset, this work first comprehensively evaluates current advanced visual object tracking methods, including SAM- and SAM2-based trackers, in challenging underwater environments, \eg, coral reefs. Our findings highlight the improvements of SAM2 over SAM, demonstrating its enhanced ability to handle the complexities of underwater camouflaged objects. Furthermore, we propose a novel vision-language tracking framework called VL-SAM2, based on the video foundation model SAM2. Extensive experimental results demonstrate that the proposed VL-SAM2 achieves state-of-the-art performance across underwater and open-air object tracking datasets. The dataset and codes are available at~{\color{magenta}{https://github.com/983632847/Awesome-Multimodal-Object-Tracking}}.

水下跟踪视觉语言SAM2伪装目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。