用视觉语言模型动态解析更新语义描述,提升目标跟踪鲁棒性
Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

- 基于视觉语言模型构建语言依赖解析机制,精准提取目标与背景关键成分
- 在TNL2K、LaSOT等四个基准上实现优于现有方法的跟踪精度与稳定性
- 适合需要高鲁棒性自然语言引导跟踪的科研与工业应用
基于自然语言规范的视觉-语言跟踪利用目标对象的高层语义线索,显著提升跟踪精度与鲁棒性。已有研究证实,在跟踪过程中自适应优化文本描述可有效缓解因目标外观、位置等属性动态变化引起的语义-视觉错配问题。然而,主流方法通过序列模型或大语言模型直接生成文本信息,不可避免存在目标误更新、背景干扰过强及广泛幻觉等问题。为此,本文提出一种新颖的语言依赖解析机制,精确提炼核心跟踪组件,包括目标对象、语义概念和背景上下文信息。在此基础上,利用预训练视觉语言模型Qwen-VL强大的跨模态理解能力,实现组件感知的自适应文本描述更新。将所设计模块集成至基线框架后,本方法在多个大规模视觉-语言跟踪基准(包括TNL2K、LaSOT、TNLLT和OTB-LANG)上均取得一致且优越的跟踪性能。源代码与预训练模型将在https://github.com/Event-AHU/Open_VLTrack发布。
原文摘要 · Abstract (English)
Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Existing studies have verified that adaptively optimizing textual descriptions throughout the tracking process can effectively mitigate the semantic-visual mismatch induced by dynamic variations in target appearance, position, and other inherent attributes. Nevertheless, mainstream methods that directly generate textual information via sequence models or large language models inevitably suffer from inherent defects, including erroneous target updating, excessive background distraction, and pervasive hallucination artifacts. To address the aforementioned limitations, this paper proposes a novel language dependency parsing mechanism to precisely distill core tracking principal components, encompassing target objects, semantic concepts, and background contextual information. On this basis, we perform component-aware adaptive textual description updates by exploiting the powerful cross-modal understanding capability of the pre-trained vision-language model Qwen-VL. By integrating the proposed elaborately designed modules into the baseline framework, our method achieves consistent and superior tracking performance on multiple large-scale vision-language tracking benchmarks, including TNL2K, LaSOT, TNLLT, and OTB-LANG. The source code and pre-trained models will be released at https://github.com/Event-AHU/Open_VLTrack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。