单图生成城市级3D场景,无需训练
Extend3D: Town-Scale 3D Generation

- 扩展隐空间实现大场景生成,分块处理并动态耦合
- 通过点云初始化与SDEdit迭代修复遮挡区域,提升完整性
- 提出'欠噪声'机制和3D感知优化,增强结构与纹理质量
本文提出Extend3D,一种无需训练的单图像3D场景生成方法,基于物体中心的3D生成模型。为突破物体中心模型在大场景中固定隐空间的限制,我们沿x、y方向扩展隐空间,并将其划分为重叠块,对每块应用物体中心3D生成模型并在每步时间耦合。由于图像与隐空间块间需严格空间对齐,我们使用单目深度估计器生成点云先验初始化场景,并通过SDEdit迭代优化遮挡区域。研究发现,在3D优化中将结构不完整视为噪声,可实现通过'欠噪声'概念完成3D补全。此外,为解决物体中心模型在子场景生成中的次优性,我们在去噪过程中优化扩展隐空间,确保去噪轨迹与子场景动态一致,并引入3D感知优化目标以提升几何结构与纹理保真度。实验表明,该方法在人类偏好与定量评估上均优于现有方法。
原文摘要 · Abstract (English)
In this paper, we propose Extend3D, a training-free pipeline for 3D scene generation from a single image, built upon an object-centric 3D generative model. To overcome the limitations of fixed-size latent spaces in object-centric models for representing wide scenes, we extend the latent space in the $x$ and $y$ directions. Then, by dividing the extended latent space into overlapping patches, we apply the object-centric 3D generative model to each patch and couple them at each time step. Since patch-wise 3D generation with image conditioning requires strict spatial alignment between image and latent patches, we initialize the scene using a point cloud prior from a monocular depth estimator and iteratively refine occluded regions through SDEdit. We discovered that treating the incompleteness of 3D structure as noise during 3D refinement enables 3D completion via a concept, which we term under-noising. Furthermore, to address the sub-optimality of object-centric models for sub-scene generation, we optimize the extended latent during denoising, ensuring that the denoising trajectories remain consistent with the sub-scene dynamics. To this end, we introduce 3D-aware optimization objectives for improved geometric structure and texture fidelity. We demonstrate that our method yields better results than prior methods, as evidenced by human preference and quantitative experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。