arXiv:2507.00454cs.CVcs.AI2025-07被引 3

解决视觉语言追踪中时空尺度不匹配问题,提升跟踪精度。

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales

  • 将语言描述按时空对应关系分解为细粒度短语并分别优化特征。
  • 引入跨帧语言令牌,增强视觉特征与语言描述的关联性。
  • 适合需要高精度多模态跟踪的场景,如智能监控与人机交互。

视觉语言追踪(VLT)的主要挑战在于目标运动导致的视觉输入与语言描述之间的错位。现有追踪器虽探索了多种有效的特征修改方法以保留更对齐的特征,但一个关键却未被充分研究的因素——视觉与语言输入在时间与空间尺度上的固有差异——仍严重制约其性能。为此,本文提出一种新型视觉语言追踪器ATSTrack,通过 extbf{A}ligning extbf{T}emporal and extbf{S}patial scale(ATST)来增强特征修改效果。具体而言,我们根据语言描述与视觉输入的时间和空间对应关系,将其分解为具有不同属性的短语,并进行细粒度特征调整。此外,引入一个包含前一帧修改后语言信息的视觉-语言令牌,引导模型提取与语言描述更相关的视觉特征,从而缓解空间尺度差异带来的影响。实验结果表明,所提ATSTrack性能可媲美现有方法。代码将公开。

原文摘要 · Abstract (English)

A main challenge of Visual-Language Tracking (VLT) is the misalignment between visual inputs and language descriptions caused by target movement. Previous trackers have explored many effective feature modification methods to preserve more aligned features. However, an important yet unexplored factor ultimately hinders their capability, which is the inherent differences in the temporal and spatial scale of information between visual and language inputs. To address this issue, we propose a novel visual-language tracker that enhances the effect of feature modification by \textbf{A}ligning \textbf{T}emporal and \textbf{S}patial scale of different input components, named as \textbf{ATSTrack}. Specifically, we decompose each language description into phrases with different attributes based on their temporal and spatial correspondence with visual inputs, and modify their features in a fine-grained manner. Moreover, we introduce a Visual-Language token that comprises modified linguistic information from the previous frame to guide the model to extract visual features that are more relevant to language description, thereby reducing the impact caused by the differences in spatial scale. Experimental results show that our proposed ATSTrack achieves performance comparable to existing methods. Our code will be released.

视觉语言目标追踪多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。