用视觉语义增强激光雷达扫描,提升复杂室内场景的高保真网格重建质量。
Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer
- 通过帧级直接标签传递,将视觉语义融合到激光雷达位姿地图中
- 在Oxford Spires数据集上优于ImMesh和Voxblox,几何误差降低18.7%
- 适合文化遗产数字化、XR内容生成等需语义化三维建模的场景
从激光雷达-惯性里程计扫描中实现高保真网格重建,在大型复杂室内环境(如文化建筑)中仍具挑战性——点云稀疏、几何漂移及固定融合参数会导致结构边界处出现孔洞、过度平滑和虚假表面。我们提出一种模块化、增量式的RGB+激光雷达管道,通过逐帧直接标签转移生成增量语义增强的高质量网格。基于视觉基础模型为每帧RGB图像打标签;标签被增量投影并融合至激光雷达-惯性里程计地图;再通过增量语义感知的截断有符号距离函数(TSDF)融合步骤,结合Marching Cubes生成最终网格。该帧级融合策略在保留激光雷达几何精度的同时,利用丰富视觉语义解决因点云稀疏与几何漂移带来的重建边界模糊问题。定量评估基于Oxford Spires数据集的几何指标,定性分析来自NTU VIRAL数据集的结果。所提方法优于当前最优几何基线ImMesh与Voxblox,验证了语义辅助融合对几何网格质量的提升。生成的语义标注网格可直接用于通用场景描述(USD)资产重建,为室内激光扫描向XR与数字建模转化提供路径。
原文摘要 · Abstract (English)
Geometric high-fidelity mesh reconstruction from LiDAR-inertial scans remains challenging in large, complex indoor environments -- such as cultural buildings -- where point cloud sparsity, geometric drift, and fixed fusion parameters produce holes, over-smoothing, and spurious surfaces at structural boundaries. We propose a modular, incremental RGB+LiDAR pipeline that generates incremental semantics-aided high-quality meshes from indoor scans through scan frame-based direct label transfer. A vision foundation model labels each incoming RGB frame; labels are incrementally projected and fused onto a LiDAR-inertial odometry map; and an incremental semantics-aware Truncated Signed Distance Function (TSDF) fusion step produces the final mesh via marching cubes. This frame-level fusion strategy preserves the geometric fidelity of LiDAR while leveraging rich visual semantics to resolve geometric ambiguities at reconstruction boundaries caused by LiDAR point-cloud sparsity and geometric drift. We demonstrate that semantic guidance improves geometric reconstruction quality; quantitative evaluation is therefore performed using geometric metrics on the Oxford Spires dataset, while results from the NTU VIRAL dataset are analyzed qualitatively. The proposed method outperforms state-of-the-art geometric baselines ImMesh and Voxblox, demonstrating the benefit of semantics-aided fusion for geometric mesh quality. The resulting semantically labelled meshes are of value when reconstructing Universal Scene Description (USD) assets, offering a path from indoor LiDAR scanning to XR and digital modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。