arXiv:2605.04574cs.CV2026-05

用视觉语言提示统一建模无人机与地面视角追踪,提升跨视角匹配可靠性。

VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

论文配图:VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
图 1 · 摘自论文原文
  • 共享编码器融合双视角特征,打破特征隔离
  • 语言提示引导跨视图交互,解决外观相似性歧义
  • 适合需要多视角协同追踪的无人机应用

无人机-地面视觉追踪(UGVT)旨在同时从无人机和地面视角追踪同一目标。现有两流方法存在特征提取孤立、依赖隐式外观匹配的问题,在视角差异大时难以建立可靠对应关系,导致追踪不可靠。为此,我们提出VL-UniTrack,一个基于视觉语言提示的完全统一框架。通过在单一共享编码器中编码双视角特征,打破特征隔离,促进充分的跨视角交互。为克服仅依赖外观匹配带来的模糊性,设计了视觉语言几何提示模块,将语言描述与视觉特征融合生成可学习提示。这些提示输入至提示引导的跨视角适配器,实现充分的跨视角特征交互,并指导视图特异性特征表示的学习。此外,提出置信度调制的互蒸馏损失,通过抑制噪声传播来正则化训练。大量实验表明,该方法在最新基准上达到最优性能。代码可在https://github.com/xuboyue1999/VL-UniTrack.git获取。

原文摘要 · Abstract (English)

UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feature extraction and rely heavily on implicit appearance matching, which struggles to establish reliable correspondence under drastic view differences, leading to tracking unreliability. To address these limitations, we propose VL-UniTrack, a fully unified framework enhanced by visual-language prompts. By encoding features from both views within a single shared encoder, our method breaks the barrier of feature isolation to facilitate sufficient cross-view interaction. To overcome the ambiguity caused by relying solely on appearance matching, we design visual-language geometric prompting module, which fuses language descriptions with visual features to generate learnable prompts. These prompts are then fed into our prompt-guided cross-view adapter module to enable sufficient cross-view feature interaction and to guide the learning of view-specific feature representations. Furthermore, a confidence-modulated mutual distillation loss is proposed to regularize the training by mitigating noise propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the latest benchmark. The code can be downloaded in https://github.com/xuboyue1999/VL-UniTrack.git

视觉追踪多视角语言提示无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。