将文本提示转为可视化热图,提升视觉语言跟踪精度
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
- 用基础模型将文本描述转为目标位置热图
- 在主流数据集上达到领先性能,显著提升跟踪准确率
- 适合需要精准语义定位的视频追踪任务
视觉语言跟踪(VLT)旨在通过视觉模板和语言描述在视频序列中定位目标。尽管文本线索能增强跟踪能力,但现有数据集图像远多于文本,导致模态对齐困难。为此,我们提出一种即插即用的新方法CTVLT,利用基础定位模型强大的图文对齐能力,将文本线索转换为可解释的视觉热图,使跟踪器更易处理。具体而言,设计文本线索映射模块,将文本描述转化为目标分布热图,直观表示文本描述的位置;同时,热图引导模块将热图与搜索图像融合,更有效地指导跟踪。在主流基准上的大量实验表明,该方法显著优于现有方法,达到当前最优性能,验证了其在增强VLT中的有效性。
原文摘要 · Abstract (English)
Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalities effectively. To address this imbalance, we propose a novel plug-and-play method named CTVLT that leverages the strong text-image alignment capabilities of foundation grounding models. CTVLT converts textual cues into interpretable visual heatmaps, which are easier for trackers to process. Specifically, we design a textual cue mapping module that transforms textual cues into target distribution heatmaps, visually representing the location described by the text. Additionally, the heatmap guidance module fuses these heatmaps with the search image to guide tracking more effectively. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our approach, achieving state-of-the-art performance and validating the utility of our method for enhanced VLT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。