arXiv:2409.19405cs.CVcs.RO2024-09ECCV被引 35

用梯度反馈让模型快速生成高质量大场景3D重建。

G3R: Gradient Guided Generalizable Reconstruction

  • 通过可微渲染的梯度信号迭代优化3D表示,结合了精调与快速预测优势。
  • 在城市驾驶和无人机数据集上提速超10倍,视觉保真度媲美甚至超越3DGS。
  • 适合需要快速高保真3D重建的自动驾驶、虚拟现实等应用。

大规模3D场景重建对虚拟现实和仿真等应用至关重要。现有神经渲染方法(如NeRF、3DGS)虽能实现大场景的逼真重建,但需逐场景优化,成本高且速度慢,且在大幅视角变化时因过拟合出现明显伪影。通用化方法或大型重建模型虽速度快,但主要适用于小场景/物体,渲染质量较低。本文提出G3R,一种可泛化的重建方法,能高效预测大场景的高质量3D表示。我们设计了一个重建网络,利用可微渲染产生的梯度反馈信号,迭代更新3D场景表示,融合了逐场景优化的高保真度与快速前馈预测的数据驱动先验。在城市驾驶和无人机数据集上的实验表明,G3R能跨多种大场景泛化,在重建速度上至少提升10倍,同时视觉真实感可媲美甚至优于3DGS,且对大幅视角变化更具鲁棒性。

原文摘要 · Abstract (English)

Large scale 3D scene reconstruction is important for applications such as virtual reality and simulation. Existing neural rendering approaches (e.g., NeRF, 3DGS) have achieved realistic reconstructions on large scenes, but optimize per scene, which is expensive and slow, and exhibit noticeable artifacts under large view changes due to overfitting. Generalizable approaches or large reconstruction models are fast, but primarily work for small scenes/objects and often produce lower quality rendering results. In this work, we introduce G3R, a generalizable reconstruction approach that can efficiently predict high-quality 3D scene representations for large scenes. We propose to learn a reconstruction network that takes the gradient feedback signals from differentiable rendering to iteratively update a 3D scene representation, combining the benefits of high photorealism from per-scene optimization with data-driven priors from fast feed-forward prediction methods. Experiments on urban-driving and drone datasets show that G3R generalizes across diverse large scenes and accelerates the reconstruction process by at least 10x while achieving comparable or better realism compared to 3DGS, and also being more robust to large view changes.

3D重建可微渲染通用化加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。