arXiv:2503.11496cs.CV2025-03被引 16

提出认知解耦框架,让机器更精准理解语言描述中的物体定位与追踪。

Cognitive Disentanglement for Referring Multi-Object Tracking

  • 模仿人类视觉系统拆分‘是什么’和‘在哪里’路径,分离语言与视觉信息
  • 在Refer-KITTI上提升6.0%的HOTA指标,显著优于现有方法
  • 适合需要高精度多目标语言追踪的自动驾驶感知场景

作为智能交通感知系统中多源信息融合的重要应用,指代式多目标跟踪(RMOT)需根据语言描述在视频序列中定位并追踪特定对象。然而,现有方法通常将语言描述视为整体嵌入,难以有效融合语言表达中丰富的语义信息与视觉特征,尤其在需同时理解静态属性与空间运动的复杂场景下表现受限。本文提出认知解耦式指代多目标跟踪(CDRMT)框架,借鉴人类视觉系统中‘什么’与‘哪里’的处理路径,先建立跨模态关联并保留模态特异性,再层次化地将语言描述解耦注入目标查询,实现从粗粒度到细粒度的语义理解优化,并基于视觉特征重构语言表示,确保追踪结果忠实反映指代表达。在多个基准数据集上的实验表明,CDRMT在Refer-KITTI上取得6.0%的平均HOTA提升,在Refer-KITTI-V2上达3.2%,显著超越当前最优方法,推动了RMOT技术的发展,并为多源信息融合提供了新视角。

原文摘要 · Abstract (English)

As a significant application of multi-source information fusion in intelligent transportation perception systems, Referring Multi-Object Tracking (RMOT) involves localizing and tracking specific objects in video sequences based on language references. However, existing RMOT approaches often treat language descriptions as holistic embeddings and struggle to effectively integrate the rich semantic information contained in language expressions with visual features. This limitation is especially apparent in complex scenes requiring comprehensive understanding of both static object attributes and spatial motion information. In this paper, we propose a Cognitive Disentanglement for Referring Multi-Object Tracking (CDRMT) framework that addresses these challenges. It adapts the "what" and "where" pathways from the human visual processing system to RMOT tasks. Specifically, our framework first establishes cross-modal connections while preserving modality-specific characteristics. It then disentangles language descriptions and hierarchically injects them into object queries, refining object understanding from coarse to fine-grained semantic levels. Finally, we reconstruct language representations based on visual features, ensuring that tracked objects faithfully reflect the referring expression. Extensive experiments on different benchmark datasets demonstrate that CDRMT achieves substantial improvements over state-of-the-art methods, with average gains of 6.0% in HOTA score on Refer-KITTI and 3.2% on Refer-KITTI-V2. Our approach advances the state-of-the-art in RMOT while simultaneously providing new insights into multi-source information fusion.

多目标跟踪语言理解视觉推理自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。