arXiv:2411.01756cs.CV2024-11NeurIPS被引 21

用大模型自动生成精准描述,提升视觉跟踪精度。

ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model

  • 通过迭代优化提示词,让大模型生成更准确的物体描述。
  • 在多个数据集上达到与顶尖方法相当的跟踪性能。
  • 适合需要高精度跟踪且有语言辅助需求的研究者。

视觉目标跟踪旨在基于初始边界框定位视频序列中的目标。近期,视觉-语言(VL)跟踪器利用自然语言描述以增强在不同场景下的泛化能力,但其跟踪性能仍落后于当前最先进的纯视觉跟踪方法。我们发现,这一差距主要源于对人工文本标注的过度依赖,尤其是频繁出现的模糊语言描述。本文提出ChatTracker,利用多模态大语言模型(MLLM)中丰富的世界知识生成高质量的语言描述,从而提升跟踪性能。为此,我们设计了一种基于反思的提示优化模块,通过跟踪反馈迭代修正目标的模糊或不准确描述。为进一步挖掘MLLM产生的语义信息,我们提出一种简单有效的VL跟踪框架,可作为即插即用模块,轻松提升各类视觉及视觉-语言跟踪器的性能。实验结果表明,所提方法在多个基准测试中达到与现有先进方法相当的性能。

原文摘要 · Abstract (English)

Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still inferior to State-of-The-Art (SoTA) visual trackers in terms of tracking performance. We found that this inferiority primarily results from their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel reflection-based prompt optimization module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed and can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that our proposed ChatTracker achieves a performance comparable to existing methods.

视觉跟踪大模型多模态语言辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。