让自动驾驶理解动态场景中带动作描述的物体指代
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
- 融合激光雷达与图像,用语言引导生成3D目标候选
- 利用运动历史和短期动态,提升指代定位准确率70%
- 适合需要理解复杂行为描述的自动驾驶系统
理解动态3D驾驶场景中的自然语言物体指代对交互式自动驾驶系统至关重要。实际中,许多指代表达通过近期运动或短时交互描述目标,仅靠静态外观或几何特征无法解决。本文研究时间性语言驱动的3D定位任务,旨在通过多帧观测识别当前帧中的被指物体。提出TrackTeller框架,统一整合激光雷达-图像融合、语言条件解码与时间推理。该框架构建与文本语义对齐的联合场景表示,生成语言感知的3D候选框,并利用运动历史与短期动态优化定位决策。在NuPrompt基准上的实验表明,TrackTeller显著提升语言引导跟踪性能,相比强基线实现平均多目标追踪准确率70%的相对提升,误报频率降低3.15至3.4倍。
原文摘要 · Abstract (English)
Understanding natural-language references to objects in dynamic 3D driving scenes is essential for interactive autonomous systems. In practice, many referring expressions describe targets through recent motion or short-term interactions, which cannot be resolved from static appearance or geometry alone. We study temporal language-based 3D grounding, where the objective is to identify the referred object in the current frame by leveraging multi-frame observations. We propose TrackTeller, a temporal multimodal grounding framework that integrates LiDAR-image fusion, language-conditioned decoding, and temporal reasoning in a unified architecture. TrackTeller constructs a shared UniScene representation aligned with textual semantics, generates language-aware 3D proposals, and refines grounding decisions using motion history and short-term dynamics. Experiments on the NuPrompt benchmark demonstrate that TrackTeller consistently improves language-grounded tracking performance, outperforming strong baselines with a 70% relative improvement in Average Multi-Object Tracking Accuracy and a 3.15-3.4 times reduction in False Alarm Frequency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。