arXiv:2411.15600cs.CVcs.AI2024-11被引 12

首次细粒度评估语言如何帮助视觉-语言追踪,揭示语义信息在不同挑战下的真实作用。

How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking

  • 构建包含10类挑战与6种语义粒度的多维评估框架
  • 在60个子空间中发现主流追踪器在复杂场景中的性能瓶颈
  • 通过解耦分析,明确不同语义类型对特定挑战的影响

视觉-语言追踪(VLT)通过引入文本信息,在快速运动、形变等复杂条件下提供语义引导,但当前VLT追踪器在多个基准上表现仍逊于单模态方法,语义信息有时反而成为干扰。为此,本文提出VLTVerse——首个细粒度评估框架,全面考虑多种挑战因素与多样语义信息,旨在揭示语言在VLT中的实际作用。贡献包括:(1) VLTVerse引入10类序列级挑战标签与6种多粒度语义信息,构建灵活且多维的评估空间;(2) 基于挑战因素与语义类型组合形成的60个子空间,系统评估三种主流SOTA VLT追踪器,揭示其在复杂场景中的性能瓶颈,并提供全新的评估视角;(3) 通过解耦分析,研究不同语义类型对特定挑战因素的影响,结合不同算法,为数据、评估与算法改进提供关键指导。VLTVerse工具包与结果将公开于 exttt{http://metaverse.aitestunion.com}。

原文摘要 · Abstract (English)

Vision-language tracking (VLT) extends traditional single object tracking by incorporating textual information, providing semantic guidance to enhance tracking performance under challenging conditions like fast motion and deformations. However, current VLT trackers often underperform compared to single-modality methods on multiple benchmarks, with semantic information sometimes becoming a "distraction." To address this, we propose VLTVerse, the first fine-grained evaluation framework for VLT trackers that comprehensively considers multiple challenge factors and diverse semantic information, hoping to reveal the role of language in VLT. Our contributions include: (1) VLTVerse introduces 10 sequence-level challenge labels and 6 types of multi-granularity semantic information, creating a flexible and multi-dimensional evaluation space for VLT; (2) leveraging 60 subspaces formed by combinations of challenge factors and semantic types, we conduct systematic fine-grained evaluations of three mainstream SOTA VLT trackers, uncovering their performance bottlenecks across complex scenarios and offering a novel perspective on VLT evaluation; (3) through decoupled analysis of experimental results, we examine the impact of various semantic types on specific challenge factors in relation to different algorithms, providing essential guidance for enhancing VLT across data, evaluation, and algorithmic dimensions. The VLTVerse, toolkit, and results will be available at \url{http://metaverse.aitestunion.com}.

视觉语言追踪细粒度评估语义引导多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。