用多传感器融合提升自动驾驶中语言引导的3D目标定位能力
Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

- 基于语言指令动态融合摄像头、激光雷达和4D雷达的物理特征
- 在新数据集上达8.05 mAP提升,单目基准上准确率超53%
- 适合研究多模态感知与智能驾驶交互的开发者
作为具身智能的关键能力,3D视觉定位(3DVG)主要在室内场景中基于RGB-D或点云输入研究,而现有室外方法大多仅依赖单目图像。这两种设置均难以满足真实户外感知需求——异构传感器捕捉互补但各异的物理特性,如视觉纹理、3D几何和物体运动,这些对灵活且鲁棒的查询自适应定位至关重要,却未被充分挖掘。为此,我们提出Talk2Sensors,首个基于相机、激光雷达和4D雷达的多传感器3D视觉定位数据集,包含8,682条语言指令和20,558个被指认目标,其提示明确对齐传感器特定物理线索。同时,我们提出TSFormer,一种统一的Transformer框架,用于自动驾驶中的语言引导3D视觉定位。该模型采用由粗到细的属性感知融合策略:语言路由属性采样器首先根据查询级语言线索调节传感器采样权重,实现粗粒度特征检索;随后稀疏保持模态仲裁器进行细粒度模态仲裁与文本引导优化,精准确定目标空间位置。该设计可根据每条提示的语义需求动态路由外观、几何与运动线索,防止密集模态压制稀疏但关键的传感信号。大量实验表明,TSFormer在多个基准上达到领先性能:在Talk2Sensors上相较最强基线提升8.05 mAP,且在单目Mono3DRefer基准上达到53.05% [email protected]。
原文摘要 · Abstract (English)
As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of real-world outdoor perception, where heterogeneous sensors capture complementary yet distinct physical properties, such as visual texture, 3D geometry, and object kinematics, that are indispensable for flexible and robust query-adaptive grounding but remain under-exploited. To bridge this gap, we introduce Talk2Sensors, the first multi-sensor 3D visual grounding dataset built upon camera, LiDAR, and 4D radar. It contains 8,682 language instructions and 20,558 referred objects, with diverse prompts explicitly aligned with sensor-specific physical cues. Furthermore, we propose TSFormer, a unified Transformer-based framework for language-guided 3D visual grounding in autonomous driving. TSFormer adopts a coarse-to-fine property-aware fusion strategy: the Language-Routed Property Sampler first performs coarse text-conditioned feature retrieval by modulating sensor sampling weights with query-level linguistic cues, while the subsequent Sparse-Preserving Modality Arbiter module conducts fine-grained modality arbitration and text-guided refinement to determine the precise referred spatial location. This design enables dynamic routing of appearance, geometry, and motion cues according to the semantic requirements of each prompt, preventing dense modalities from overwhelming sparse but critical sensor signals. Extensive experiments demonstrate that TSFormer achieves state-of-the-art performance across multiple benchmarks: it improves over the strongest baseline by 8.05 mAP on Talk2Sensors, and transfers to the monocular Mono3DRefer benchmark with 53.05\% [email protected].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。