arXiv:2503.18035cs.CV2025-03被引 1

用文本驱动3D点云定位,提升自动驾驶环境识别精度。

Vehicle-Scene Interaction: A Text-Driven 3D Lidar Place Recognition Method for Autonomous Driving

  • 分两阶段:先融合多尺度几何特征,再通过跨模态对齐增强文本与点云匹配
  • 在KITTI360上达40%顶1准确率,比现有方法高7个百分点
  • 适合需要文本控制的智能车、配送机器人等大场景定位任务

基于环境描述的大规模点云地图中的定位,在如最后一公里配送机器人等大规模自主系统发展中至关重要。然而,当前方法因点云编码器难以有效捕捉局部细节和长程空间关系,以及文本与点云表征间存在显著模态差异而面临挑战。为此,我们提出Des4Pos,一种新颖的两阶段文本驱动遥感定位框架。粗粒度阶段,点云编码器采用多尺度融合注意力机制(MFAM)增强局部几何特征,并通过双向LSTM模块强化全局空间关系;同时,步进式文本编码器(STE)融合CLIP[1]提供的跨模态先验知识,利用该先验对齐文本与点云特征,有效弥合模态差距。细粒度阶段引入级联残差注意力(CRA)模块融合跨模态特征并预测相对定位偏移,从而实现更高精度定位。在KITTI360Pose测试集上的实验表明,Des4Pos在文本到点云的场景识别中达到最先进性能,具体在5米半径阈值下,顶1准确率达40%,顶10准确率达77%,分别优于最佳现有方法7%和7%。

原文摘要 · Abstract (English)

Environment description-based localization in large-scale point cloud maps constructed through remote sensing is critically significant for the advancement of large-scale autonomous systems, such as delivery robots operating in the last mile. However, current approaches encounter challenges due to the inability of point cloud encoders to effectively capture local details and long-range spatial relationships, as well as a significant modality gap between text and point cloud representations. To address these challenges, we present Des4Pos, a novel two-stage text-driven remote sensing localization framework. In the coarse stage, the point-cloud encoder utilizes the Multi-scale Fusion Attention Mechanism (MFAM) to enhance local geometric features, followed by a bidirectional Long Short-Term Memory (LSTM) module to strengthen global spatial relationships. Concurrently, the Stepped Text Encoder (STE) integrates cross-modal prior knowledge from CLIP [1] and aligns text and point-cloud features using this prior knowledge, effectively bridging modality discrepancies. In the fine stage, we introduce a Cascaded Residual Attention (CRA) module to fuse cross-modal features and predict relative localization offsets, thereby achieving greater localization precision. Experiments on the KITTI360Pose test set demonstrate that Des4Pos achieves state-of-the-art performance in text-to-point-cloud place recognition. Specifically, it attains a top-1 accuracy of 40% and a top-10 accuracy of 77% under a 5-meter radius threshold, surpassing the best existing methods by 7% and 7%, respectively.

3D定位文本驱动点云匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。