arXiv:2508.13228cs.GRcs.AI2025-08

用多模态信息加速高精度3D场景重建,训练更快更准。

PreSem-Surf: RGB-D Surface Reconstruction with Progressive Semantic Modeling and SG-MLP Pre-Rendering Mechanism

  • 结合图像、深度和语义信息,分阶段提升建模精度。
  • 在7个合成场景上,3项指标领先,其他指标保持竞争力。
  • 适合需要快速高质量3D重建的工业与科研应用。

本文提出PreSem-Surf,一种基于神经辐射场(NeRF)框架的优化方法,可从RGB-D序列中快速重建高质量场景表面。该方法融合图像、深度与语义信息以提升重建性能。具体而言,引入新型SG-MLP采样结构与PR-MLP(预条件多层感知机)实现体素预渲染,使模型更早捕捉场景信息并更好区分噪声与局部细节。同时采用渐进式语义建模,在逐步提高精度的同时减少训练时间,增强场景理解能力。在七个合成场景上,使用六项评估指标进行实验,PreSem-Surf在C-L1、F-score和IoU三项指标上表现最佳,且在NC、Accuracy和Completeness上保持竞争力,验证了其有效性和实用性。

原文摘要 · Abstract (English)

This paper proposes PreSem-Surf, an optimized method based on the Neural Radiance Field (NeRF) framework, capable of reconstructing high-quality scene surfaces from RGB-D sequences in a short time. The method integrates RGB, depth, and semantic information to improve reconstruction performance. Specifically, a novel SG-MLP sampling structure combined with PR-MLP (Preconditioning Multilayer Perceptron) is introduced for voxel pre-rendering, allowing the model to capture scene-related information earlier and better distinguish noise from local details. Furthermore, progressive semantic modeling is adopted to extract semantic information at increasing levels of precision, reducing training time while enhancing scene understanding. Experiments on seven synthetic scenes with six evaluation metrics show that PreSem-Surf achieves the best performance in C-L1, F-score, and IoU, while maintaining competitive results in NC, Accuracy, and Completeness, demonstrating its effectiveness and practical applicability.

3D重建多模态融合神经渲染高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。