NeVStereo统一实现高精度3D重建,从随意拍摄图像中同时输出相机位姿、深度图、新视角合成与表面重建。
NeVStereo: A NeRF-Driven NVS-Stereo Architecture for High-Fidelity 3D Tasks
- 基于NeRF的立体友好渲染与置信度引导的多视图深度估计
- 通过迭代优化提升深度与辐射场的一致性,降低表面堆叠和伪影
- 零样本测试下深度误差降36%,新视角合成质量超主流方法4.5%
在现代密集3D重建中,前馈系统(如VGGT、pi3)侧重端到端匹配与几何预测,但未显式输出新视角合成(NVS)。基于神经渲染的方法能提供高保真NVS与详细几何,但通常假设相机位姿固定,对位姿误差敏感。因此,如何从随意拍摄的多视图图像中同时获得准确位姿、可靠深度、高质量渲染和精确3D表面仍具挑战。我们提出NeVStereo,一种基于NeRF的NVS-立体架构,可联合生成相机位姿、多视图深度、新视角合成与表面重建,仅需多视图RGB输入。其结合了NeRF驱动的立体友好渲染、置信度引导的多视图深度估计、耦合束调整的位姿优化及迭代更新深度与辐射场的阶段,有效缓解了常见NeRF问题如表面堆叠、伪影与位姿-深度耦合。在室内、室外、桌面与航拍基准上,实验显示该方法在零样本场景下表现稳定,深度误差降低最多达36%,位姿精度提升10.4%,新视角合成质量提高4.5%,网格质量达到当前最优(F1 91.93%,Chamfer 4.35 mm)。
原文摘要 · Abstract (English)
In modern dense 3D reconstruction, feed-forward systems (e.g., VGGT, pi3) focus on end-to-end matching and geometry prediction but do not explicitly output the novel view synthesis (NVS). Neural rendering-based approaches offer high-fidelity NVS and detailed geometry from posed images, yet they typically assume fixed camera poses and can be sensitive to pose errors. As a result, it remains non-trivial to obtain a single framework that can offer accurate poses, reliable depth, high-quality rendering, and accurate 3D surfaces from casually captured views. We present NeVStereo, a NeRF-driven NVS-stereo architecture that aims to jointly deliver camera poses, multi-view depth, novel view synthesis, and surface reconstruction from multi-view RGB-only inputs. NeVStereo combines NeRF-based NVS for stereo-friendly renderings, confidence-guided multi-view depth estimation, NeRF-coupled bundle adjustment for pose refinement, and an iterative refinement stage that updates both depth and the radiance field to improve geometric consistency. This design mitigated the common NeRF-based issues such as surface stacking, artifacts, and pose-depth coupling. Across indoor, outdoor, tabletop, and aerial benchmarks, our experiments indicate that NeVStereo achieves consistently strong zero-shot performance, with up to 36% lower depth error, 10.4% improved pose accuracy, 4.5% higher NVS fidelity, and state-of-the-art mesh quality (F1 91.93%, Chamfer 4.35 mm) compared to existing prestigious methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。