arXiv:2503.08336cs.CV2025-03

融合激光雷达与毫米波雷达点云,提升自动驾驶3D视觉定位精度

Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving

  • 采用双阶段异构模态自适应融合,结合激光雷达与雷达数据
  • 在Talk2Radar和Talk2Car数据集上达到当前最优性能
  • 适合关注多传感器融合与3D视觉理解的研究者

具身室外场景理解是自主智能体感知、分析并响应动态驾驶环境的基础。现有3D理解主要依赖2D视觉-语言模型(VLMs),但其场景上下文信息有限。相比之下,激光雷达(LiDAR)提供丰富的深度与精细3D物体表征,而新兴的4D毫米波雷达能检测每个目标的运动趋势、速度和反射强度。两者的融合为自然语言查询提供了更灵活的条件,从而支持更精确的3D视觉定位。为此,我们提出首个基于提示引导点云传感器融合范式的室外3D视觉定位模型TPCNet。为优化提示所需的双传感器特征融合,设计了两阶段异构模态自适应融合:首先使用双向代理交叉注意力(BACA)将双传感器特征与文本特征对齐;其次引入动态门控图融合(DGGF)模块定位查询区域。为进一步提升精度,提出基于最近物体边缘的C3D-RECHead。实验表明,TPCNet及其各模块在Talk2Radar和Talk2Car数据集上均达到最先进水平。代码已开源。

原文摘要 · Abstract (English)

Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. However, existing 3D understanding is predominantly based on 2D Vision-Language Models (VLMs), which collect and process limited scene-aware contexts. In contrast, compared to the 2D planar visual information, point cloud sensors such as LiDAR provide rich depth and fine-grained 3D representations of objects. Even better the emerging 4D millimeter-wave radar detects the motion trend, velocity, and reflection intensity of each object. The integration of these two modalities provides more flexible querying conditions for natural language, thereby supporting more accurate 3D visual grounding. To this end, we propose a novel method called TPCNet, the first outdoor 3D visual grounding model upon the paradigm of prompt-guided point cloud sensor combination, including both LiDAR and radar sensors. To optimally combine the features of these two sensors required by the prompt, we design a multi-fusion paradigm called Two-Stage Heterogeneous Modal Adaptive Fusion. Specifically, this paradigm initially employs Bidirectional Agent Cross-Attention (BACA), which feeds both-sensor features, characterized by global receptive fields, to the text features for querying. Moreover, we design a Dynamic Gated Graph Fusion (DGGF) module to locate the regions of interest identified by the queries. To further enhance accuracy, we devise an C3D-RECHead, based on the nearest object edge to the ego-vehicle. Experimental results demonstrate that our TPCNet, along with its individual modules, achieves the state-of-the-art performance on both the Talk2Radar and Talk2Car datasets. We release the code at https://github.com/GuanRunwei/TPCNet.

3D视觉定位多传感器融合自动驾驶点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。