arXiv:2511.01186cs.ROcs.CV2025-11被引 7

融合激光雷达与视觉模型,实现大场景高精度点云重建

LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping

  • 分两阶段融合激光雷达与视觉信息,先粗后细提升定位精度
  • 在多个数据集上实现全局一致且带度量尺度的稠密彩色点云
  • 适合需要高精度三维建图的机器人与自动驾驶场景

构建大规模彩色点云是机器人感知、导航与场景理解的重要任务。尽管激光雷达惯性视觉里程计(LIVO)取得进展,其性能仍高度依赖外部标定。同时,如VGGT等3D视觉基础模型在大环境中的可扩展性有限,且缺乏度量尺度。为此,我们提出LiDAR-VGGT,通过两级粗到细的跨模态融合管道,将LIVO与最先进的VGGT模型紧密结合:首先,预融合模块通过鲁棒初始化优化,在每个会话中高效估计具有粗略度量尺度的VGGT位姿与点云;随后,后融合模块利用基于边界框的正则化,增强跨模态3D相似性变换,减少因激光雷达与相机视场不一致导致的尺度失真。多数据集实验证明,LiDAR-VGGT生成了稠密、全局一致且带度量尺度的彩色点云,显著优于基于VGGT的方法和LIVO基线。本文提出的新型彩色点云评估工具包将开源。

原文摘要 · Abstract (English)

Reconstructing large-scale colored point clouds is an important task in robotics, supporting perception, navigation, and scene understanding. Despite advances in LiDAR inertial visual odometry (LIVO), its performance remains highly sensitive to extrinsic calibration. Meanwhile, 3D vision foundation models, such as VGGT, suffer from limited scalability in large environments and inherently lack metric scale. To overcome these limitations, we propose LiDAR-VGGT, a novel framework that tightly couples LiDAR inertial odometry with the state-of-the-art VGGT model through a two-stage coarse- to-fine fusion pipeline: First, a pre-fusion module with robust initialization refinement efficiently estimates VGGT poses and point clouds with coarse metric scale within each session. Then, a post-fusion module enhances cross-modal 3D similarity transformation, using bounding-box-based regularization to reduce scale distortions caused by inconsistent FOVs between LiDAR and camera sensors. Extensive experiments across multiple datasets demonstrate that LiDAR-VGGT achieves dense, globally consistent colored point clouds and outperforms both VGGT-based methods and LIVO baselines. The implementation of our proposed novel color point cloud evaluation toolkit will be released as open source.

三维重建多模态融合点云建图自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。