arXiv:2410.02492cs.CVcs.CL2024-10被引 17

用大模型生成多样文本,构建更真实的视觉语言追踪测试集

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM

  • 用大模型生成不同长度和粒度的语义描述,提升文本多样性
  • 基于5个主流数据集构建新基准,包含短时、长时、全局追踪三任务
  • 揭示现有算法在多样文本下的性能瓶颈,助力视频理解研究

视觉语言追踪(VLT)作为前沿研究方向,利用语言信息增强多模态输入,将传统单目标追踪(SOT)扩展至视频理解应用。然而,现有大多数VLT基准仍依赖简洁的人工标注文本,难以捕捉视频内容动态细节,且语言风格单一、描述粒度固定,导致算法倾向于采用‘记忆答案’策略,偏离深层理解目标。本文利用大语言模型(LLM)生成多样化语义标注(涵盖文本长度与粒度差异),构建新型多模态基准DTVLT。具体包括:(1) 基于五个代表性VLT与SOT基准,构建含短时、长时、全局实例追踪三子任务的新基准DTVLT;(2) 提出四种粒度文本,覆盖语义信息的广度与密度,促进算法对复杂视频的理解能力;(3) 在DTVLT上开展全面实验,分析多样文本对追踪性能的影响,并揭示现有算法的性能瓶颈。相关基准、实验结果与工具包将陆续发布于http://videocube.aitestunion.com/。

原文摘要 · Abstract (English)

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to encompass video understanding applications. Despite this, most VLT benchmarks still depend on succinct, human-annotated text descriptions for each video. These descriptions often fall short in capturing the nuances of video content dynamics and lack stylistic variety in language, constrained by their uniform level of detail and a fixed annotation frequency. As a result, algorithms tend to default to a "memorize the answer" strategy, diverging from the core objective of achieving a deeper understanding of video content. Fortunately, the emergence of large language models (LLMs) has enabled the generation of diverse text. This work utilizes LLMs to generate varied semantic annotations (in terms of text lengths and granularities) for representative SOT benchmarks, thereby establishing a novel multi-modal benchmark. Specifically, we (1) propose a new visual language tracking benchmark with diverse texts, named DTVLT, based on five prominent VLT and SOT benchmarks, including three sub-tasks: short-term tracking, long-term tracking, and global instance tracking. (2) We offer four granularity texts in our benchmark, considering the extent and density of semantic information. We expect this multi-granular generation strategy to foster a favorable environment for VLT and video understanding research. (3) We conduct comprehensive experimental analyses on DTVLT, evaluating the impact of diverse text on tracking performance and hope the identified performance bottlenecks of existing algorithms can support further research in VLT and video understanding. The proposed benchmark, experimental results and toolkit will be released gradually on http://videocube.aitestunion.com/.

视觉追踪多模态大模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。