提出融合图像、深度与语言的多目标跟踪新任务,提升复杂场景下的定位与追踪精度。
DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking

- 构建多模态融合的深度感知跟踪框架,利用图像、深度和语言联合建模。
- 在187个场景的DRSet数据集上实现56条含深度描述的精准跟踪,显著提升遮挡下身份一致性。
- 适合研究视觉-语言-空间联合建模、机器人交互与自动驾驶中的目标追踪方向。
参照式多目标跟踪(RMOT)旨在基于语言描述追踪特定目标,对机器人和自动驾驶等交互式AI系统至关重要。然而现有模型仅依赖2D RGB数据,在处理涉及复杂空间语义(如“离摄像头最近的人”)或严重遮挡时难以准确检测与关联目标。本文提出新任务——RGBD参照式多目标跟踪(DRMOT),要求模型融合RGB、深度(D)与语言(L)三模态信息实现3D感知跟踪。为此,我们构建了专用数据集DRSet,包含187个场景的RGB图像与深度图,以及240条语言描述,其中56条包含深度相关信息。同时提出DRTrack框架,通过多模态大模型引导的深度感知目标定位与深度线索增强的轨迹关联,实现鲁棒跟踪。在DRSet上的大量实验验证了该框架的有效性。
原文摘要 · Abstract (English)
Referring Multi-Object Tracking (RMOT) aims to track specific targets based on language descriptions and is vital for interactive AI systems such as robotics and autonomous driving. However, existing RMOT models rely solely on 2D RGB data, making it challenging to accurately detect and associate targets characterized by complex spatial semantics (e.g., ``the person closest to the camera'') and to maintain reliable identities under severe occlusion, due to the absence of explicit 3D spatial information. In this work, we propose a novel task, RGBD Referring Multi-Object Tracking (DRMOT), which explicitly requires models to fuse RGB, Depth (D), and Language (L) modalities to achieve 3D-aware tracking. To advance research on the DRMOT task, we construct a tailored RGBD referring multi-object tracking dataset, named DRSet, designed to evaluate models' spatial-semantic grounding and tracking capabilities. Specifically, DRSet contains RGB images and depth maps from 187 scenes, along with 240 language descriptions, among which 56 descriptions incorporate depth-related information. Furthermore, we propose DRTrack, a MLLM-guided depth-referring tracking framework. DRTrack performs depth-aware target grounding from joint RGB-D-L inputs and enforces robust trajectory association by incorporating depth cues. Extensive experiments on the DRSet dataset demonstrate the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。