arXiv:2505.16321cs.CV2025-05ICML被引 5

用轻量运动提示提升视觉追踪鲁棒性,不增加速度负担

Efficient Motion Prompt Learning for Robust Visual Tracking

  • 引入运动提示模块,融合长期运动轨迹与视觉特征
  • 在7个基准上显著提升追踪鲁棒性,训练成本低、速度几乎不变
  • 可无缝接入现有追踪器,适合追求稳定性与效率的开发者

由于处理时序信息的挑战,多数追踪器仅依赖视觉区分性,忽视视频数据的独特时序一致性。本文提出一种轻量级、即插即用的运动提示追踪方法,可轻松集成至现有视觉追踪器中,构建联合追踪框架,同时利用运动与视觉线索,实现高效提示学习下的鲁棒追踪。设计了三种位置编码的运动编码器,将长期运动轨迹编码至视觉嵌入空间;并引入融合解码器与自适应权重机制,动态融合视觉与运动特征。将该运动模块集成至三个不同追踪器,共五种模型。在七个挑战性追踪基准上的实验表明,所提模块显著提升视觉追踪器的鲁棒性,训练成本极低,速度几乎无损失。代码已开源。

原文摘要 · Abstract (English)

Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking.

视觉追踪运动提示轻量模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。