arXiv:2409.08887cs.CVcs.CL2024-09被引 17

首个支持多轮交互的视觉语言追踪基准,让跟踪更贴近真实人机协作。

Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark

  • 引入多轮文本与目标框更新机制,实现持续交互式追踪
  • 在失败时通过修正文本和框提升追踪准确率与鲁棒性
  • 为多模态跟踪器提供细粒度评估新范式,适合研究交互式视觉理解者

视觉语言追踪(VLT)通过融合语言语义信息,克服仅依赖视觉模态的局限,推动更高级的人机交互。其核心是认知对齐,通常需多轮信息交互,尤其在序列决策过程中。然而现有VLT基准仅在首帧提供初始文本和边界框,后续无交互,违背了原任务初衷。为此,本文提出首个引入多轮交互的稳健基准VLT-MI(Visual Language Tracking with Multi-modal Interaction)。首先,基于主流基准,利用DTLLM-VLT生成多样、多层次的多轮交互文本,借助大模型的世界知识;其次,提出新交互范式,通过文本更新与目标恢复实现多轮交互;当出现多次追踪失败时,通过提供更对齐的文本与修正后的边界框,扩展下游任务能力。在传统VLT基准与VLT-MI上开展对比实验,评估并分析跟踪器在交互范式下的精度与鲁棒性。本工作为VLT任务提供新视角与范式,支持多模态跟踪器的精细评估。未来可拓展至更多数据集,促进视频-语言模型能力的广泛评测。

原文摘要 · Abstract (English)

Visual Language Tracking (VLT) enhances tracking by mitigating the limitations of relying solely on the visual modality, utilizing high-level semantic information through language. This integration of the language enables more advanced human-machine interaction. The essence of interaction is cognitive alignment, which typically requires multiple information exchanges, especially in the sequential decision-making process of VLT. However, current VLT benchmarks do not account for multi-round interactions during tracking. They provide only an initial text and bounding box (bbox) in the first frame, with no further interaction as tracking progresses, deviating from the original motivation of the VLT task. To address these limitations, we propose a novel and robust benchmark, VLT-MI (Visual Language Tracking with Multi-modal Interaction), which introduces multi-round interaction into the VLT task for the first time. (1) We generate diverse, multi-granularity texts for multi-round, multi-modal interaction based on existing mainstream VLT benchmarks using DTLLM-VLT, leveraging the world knowledge of LLMs. (2) We propose a new VLT interaction paradigm that achieves multi-round interaction through text updates and object recovery. When multiple tracking failures occur, we provide the tracker with more aligned texts and corrected bboxes through interaction, thereby expanding the scope of VLT downstream tasks. (3) We conduct comparative experiments on both traditional VLT benchmarks and VLT-MI, evaluating and analyzing the accuracy and robustness of trackers under the interactive paradigm. This work offers new insights and paradigms for the VLT task, enabling a fine-grained evaluation of multi-modal trackers. We believe this approach can be extended to additional datasets in the future, supporting broader evaluations and comparisons of video-language model capabilities.

视觉追踪多模态交互语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。