arXiv:2502.02657cs.ROcs.CV2025-02中稿 · T-RO被引 5

融合激光雷达与视觉数据,实现大场景高精度辐射场重建并量化不确定性。

SiLVR: Scalable Lidar-Visual Radiance Field Reconstruction with Uncertainty Quantification

  • 基于NeRF框架融合激光雷达与视觉数据,提升几何精度。
  • 通过空间方差估计实现不确定性量化,识别低可信度区域。
  • 适用于机器人和手持扫描,覆盖面积超2万平米,适合高精度重建需求。

我们提出一种基于神经辐射场(NeRF)的大规模重建系统,融合激光雷达与视觉数据,生成几何精确且纹理逼真的三维重建结果。该系统采用先进NeRF表示,引入激光雷达数据以提供强几何约束,尤其在纹理单一、视觉线索模糊的区域表现更优。核心贡献是提出一种新方法,通过估计每个点位置在观测数据下的空间方差,实现对激光雷达-视觉NeRF重建的后验认知不确定性量化。这为评估各传感器模态贡献提供了理论依据,可识别因均匀纹理、视角受限或激光雷达覆盖不足导致的不确定重建并移除。系统集成实时激光雷达SLAM,用于启动结构光运动(SfM)重建,并正确约束整体度量尺度,这对激光雷达深度损失至关重要。优化后的SLAM轨迹使用谱聚类划分为子地图,按共视性分组图像,优于距离分割。该子地图划分在拼接时尤为有效,因边界常受观测限制而产生伪影。实验在多相机-激光雷达系统上进行,涵盖机器人搭载与手持扫描场景,测试数据集总面积超过20,000平方米。

原文摘要 · Abstract (English)

We present a neural radiance field (NeRF) based large-scale reconstruction system that fuses lidar and vision data to generate high-quality reconstructions that are geometrically accurate and capture photorealistic texture. Our system adopts the state-of-the-art NeRF representation to incorporate lidar. Adding lidar data adds strong geometric constraints on the depth and surface normals, which is particularly useful when modelling uniform texture surfaces which contain ambiguous visual reconstruction cues. A key contribution of this work is a novel method to quantify the epistemic uncertainty of the lidar-visual NeRF reconstruction by estimating the spatial variance of each point location in the radiance field given the sensor observations from the cameras and lidar. This provides a principled approach to evaluate the contribution of each sensor modality to the final reconstruction. In this way, reconstructions that are uncertain (due to e.g. uniform visual texture, limited observation viewpoints, or little lidar coverage) can be identified and removed. Our system is integrated with a real-time lidar SLAM system which is used to bootstrap a Structure-from-Motion (SfM) reconstruction procedure. It also helps to properly constrain the overall metric scale which is essential for the lidar depth loss. The refined SLAM trajectory can then be divided into submaps using Spectral Clustering to group sets of co-visible images together. This submapping approach is more suitable for visual reconstruction than distance-based partitioning. Our uncertainty estimation is particularly effective when merging submaps as their boundaries often contain artefacts due to limited observations. We demonstrate the reconstruction system using a multi-camera, lidar sensor suite in experiments involving both robot-mounted and handheld scanning. Our test datasets cover a total area of more than 20,000 square metres.

NeRF激光雷达三维重建不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。