arXiv:2605.23888cs.CV2026-05被引 1

用生成模型提升多视角3D场景重建质量,精度比顶尖方法高16%。

GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

论文配图:GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction
图 1 · 摘自论文原文
  • 将3D场景拆分为重叠块,结合生成式先验进行条件生成
  • 在室内场景上实现16%的精度提升,生成可编辑的PBR网格
  • 不依赖视角顺序,能保持多视角一致性,适合高质量重建

我们提出一种从多视角RGB图像进行高保真3D场景重建的新方法,通过紧密耦合强大的生成式3D先验来实现。将场景重建建模为对一系列空间局部、重叠的块进行条件3D生成,从而扩展生成能力至大场景范围。关键在于继承最先进的生成形状模型(以Trellis.2为例)的保真度与完整性,并将其推广至场景层面。为此,我们设计了一种基于投影的条件机制,将带有姿态的多视角图像特征映射到与生成模型对齐的连贯3D表示中,该表示独立于视角顺序且空间锚定于场景,实现了高保真、多视角一致的生成几何。这使得Trellis.2在物体级别的强先验得以应用于多视角、场景级生成,生成了忠实且可编辑的室内环境PBR网格。最终结果在高保真度上优于当前顶尖重建方法16%。

原文摘要 · Abstract (English)

We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D generation over a set of spatially-localized, overlapping chunks that together tile the scene, scaling generation to large scene extents. Crucially, we inherit the fidelity and completeness of state-of-the-art generative shape models -- we use Trellis.2 as an example -- which we generalize to the scene level. To this end, we propose a projection-based conditioning mechanism that lifts posed multi-view image features into a coherent 3D representation aligned with the generative model, independent of view ordering and spatially anchored to the scene, yielding high-fidelity, multi-view consistent generated geometry. This enables lifting the strong object-level prior of Trellis.2 to multi-view, scene-scale generation, producing faithful, editable PBR mesh reconstructions of indoor environments. As a result, we obtain high-fidelity results that outperform cutting-edge reconstruction methods by 16%.

3D重建生成模型多视角场景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。