用单图重建户外场景,细节更准结构更稳
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
- 融合深度与法向信息,提升几何细节捕捉能力
- 提出几何仿射场与3D自注意力,兼顾结构一致与渲染效率
- 基于深度的3D高斯解码器,支持高质量新视角生成
单视角3D场景重建在高保真室外场景建模中仍面临细节丢失与结构不一致的挑战。本文提出Niagara框架,首次实现从单张图像高保真重建复杂室外场景。该方法融合单目深度与法向估计作为输入,显著提升对细微几何结构的还原能力,缓解常见失真问题。引入几何仿射场(GAF)与3D自注意力机制作为几何约束,结合显式几何的结构特性与隐式特征场的适应性,平衡高效渲染与高精度重建。框架采用专用编码器-解码器结构,提出基于深度的3D高斯解码器,用于预测3D高斯参数以支持新视角合成。大量实验表明,Niagara在单视图与双视图设置下均超越Flash3D等现有最先进方法,尤其在室外场景中显著提升几何准确度与视觉保真度。
原文摘要 · Abstract (English)
Recent advances in single-view 3D scene reconstruction have highlighted the challenges in capturing fine geometric details and ensuring structural consistency, particularly in high-fidelity outdoor scene modeling. This paper presents Niagara, a new single-view 3D scene reconstruction framework that can faithfully reconstruct challenging outdoor scenes from a single input image for the first time. Our approach integrates monocular depth and normal estimation as input, which substantially improves its ability to capture fine details, mitigating common issues like geometric detail loss and deformation. Additionally, we introduce a geometric affine field (GAF) and 3D self-attention as geometry-constraint, which combines the structural properties of explicit geometry with the adaptability of implicit feature fields, striking a balance between efficient rendering and high-fidelity reconstruction. Our framework finally proposes a specialized encoder-decoder architecture, where a depth-based 3D Gaussian decoder is proposed to predict 3D Gaussian parameters, which can be used for novel view synthesis. Extensive results and analyses suggest that our Niagara surpasses prior SoTA approaches such as Flash3D in both single-view and dual-view settings, significantly enhancing the geometric accuracy and visual fidelity, especially in outdoor scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。