用自适应文本注入提升语言引导的单目标追踪精度
Efficient Language-to-Vision Feature Injection for Referring Single-Object Tracking

- 通过门控特征注入机制动态调节语言信息强度
- 在多个基准上达到领先性能,无需视觉语言对齐训练
- 适合需要低训练成本的语言-视觉追踪场景
语言引导的单目标追踪通过联合利用语义线索和视觉模板实现目标初始化与后续追踪。核心挑战在于语言信息在不同阶段应差异化使用:初始时不可或缺,但追踪过程中过度强调会导致语义漂移。现有方法通常需昂贵的视觉-语言对齐训练。本文提出LVTrack,一种纯Transformer框架,引入模式感知的门控特征注入器,自适应调节文本引导,缓解语义漂移。结合针对性适配,直接使用冻结的视觉-语言预训练模型,大幅降低训练成本并保持强语言理解能力。为提升时间定位效果,LVTrack融合混合相对-绝对位置编码与轻量记忆机制,并采用高斯平滑KL损失优化自回归框预测。在标准基准上的大量实验表明,该方法表现优异。
原文摘要 · Abstract (English)
Referring single-object tracking enables language-grounded target initialization and subsequent tracking by jointly leveraging semantic cues and visual templates. The core difficulty is to use language differently across stages: it is indispensable for grounding but can induce semantic drift during tracking when overemphasized. Meanwhile, current methods often require costly vision-language alignment training. We present LVTrack, a pure transformer framework that introduces a mode-conditioned Gated Feature Injector to adaptively regulate textual guidance and alleviate semantic drift. Together with targeted adaptations, it directly harnesses a frozen vision-language pretrained model, greatly reducing training cost and preserving strong language understanding. To further improve temporal localization, LVTrack integrates hybrid relative-absolute positional encodings with a lightweight memory mechanism and optimizes autoregressive box prediction using a Gaussian-smoothed KL loss. Extensive experiments on standard benchmarks demonstrate that LVTrack achieves strong performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。