arXiv:2511.17681cs.CV2025-11中稿 · IEEE Transactions …被引 3

让语言描述跟上物体运动变化,提升多目标跟踪精度。

Motion-Aware Vision-Reference Alignment for Referring Multi-Object Tracking

  • 用物体动态行为生成运动模态,与视觉和语言对齐
  • 在多个基准上超越现有最先进方法,显著提升追踪准确率
  • 适合需要精准动态理解的多模态跟踪场景

指代式多目标跟踪(RMOT)通过自然语言实现多模态融合跟踪,但现有基准仅描述物体外观、相对位置和初始运动状态,忽略运动过程中的速度变化与方向改变。这种静态描述与动态视觉信息存在时间错配,制约了多模态跟踪性能。为此,我们提出首个面向运动感知的视觉-语言对齐框架VMRMOT。该框架从物体动态行为中提取运动模态,利用多模态大模型(MLLMs)生成运动描述并编码为运动特征。设计层次化视觉-运动-语言对齐模块(VMRA),增强跨模态一致性;同时引入运动引导预测头(MGPH),利用运动信息提升预测性能。在多个RMOT基准上的实验证明,VMRMOT显著优于现有方法。代码已开源。

原文摘要 · Abstract (English)

Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, relative positions, and initial motion states. This so-called static regulation fails to capture dynamic changes of the object motion, including velocity changes and motion direction shifts. This limitation not only causes a temporal discrepancy between static references and dynamic vision modality but also constrains multi-modal tracking performance. To address this limitation, we propose a novel motion-aware vision-reference alignment framework, named VMRMOT. VMRMOT introduces a motion modality derived from object dynamics to facilitate the alignment between the vision modality and language references. Specifically, we introduce motion-aware descriptions derived from object dynamic behaviors and encode them into motion features as the motion modality through multi-modal large language models (MLLMs). We further design a Vision-Motion-Reference Alignment (VMRA) module to hierarchically align visual queries with motion and reference cues, enhancing their cross-modal consistency. In addition, a Motion-Guided Prediction Head (MGPH) is developed to explore motion modality to enhance the performance of the prediction head. To the best of our knowledge, VMRMOT is the first motion-aware vision-reference alignment framework for the RMOT task. Extensive experiments on multiple RMOT benchmarks demonstrate that VMRMOT outperforms existing state-of-the-art methods. The code is available at https://github.com/Kroery/VMRMOT.

多目标跟踪多模态对齐运动建模语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。