用布局引导生成复杂室内3D场景,支持多房间不规则结构。
SceneCraft: Layout-Guided 3D Scene Generation

- 通过渲染生成多视角2D代理图,将语义布局转为图像输入。
- 结合语义与深度条件扩散模型生成图像,驱动神经辐射场重建场景。
- 可生成多卧室公寓级复杂场景,视觉真实且几何一致。
传统3D建模工具在创建符合用户需求的复杂3D场景时效率低下。尽管已有文本到3D生成方法,但普遍局限于小规模场景,对形状与纹理控制有限。本文提出SceneCraft,一种基于布局引导的详细室内场景生成方法,支持用户提供的文本描述与空间布局偏好。核心是渲染驱动技术,将3D语义布局转化为多视角2D代理图;并设计语义与深度条件扩散模型生成多视角图像,用于训练神经辐射场(NeRF)作为最终场景表示。不受全景图像生成限制,本方法可生成超越单间、包含不规则结构的复杂室内空间,如整套多卧室公寓。实验表明,相比现有方法,SceneCraft在复杂室内场景生成中展现出更丰富纹理、一致几何结构与逼真视觉质量。代码与更多结果见:https://orangesodahub.github.io/SceneCraft
原文摘要 · Abstract (English)
The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have achieved automatic text-to-3D generation, they are generally limited to small-scale scenes with restricted control over the shape and texture. We introduce SceneCraft, a novel method for generating detailed indoor scenes that adhere to textual descriptions and spatial layout preferences provided by users. Central to our method is a rendering-based technique, which converts 3D semantic layouts into multi-view 2D proxy maps. Furthermore, we design a semantic and depth conditioned diffusion model to generate multi-view images, which are used to learn a neural radiance field (NeRF) as the final scene representation. Without the constraints of panorama image generation, we surpass previous methods in supporting complicated indoor space generation beyond a single room, even as complicated as a whole multi-bedroom apartment with irregular shapes and layouts. Through experimental analysis, we demonstrate that our method significantly outperforms existing approaches in complex indoor scene generation with diverse textures, consistent geometry, and realistic visual quality. Code and more results are available at: https://orangesodahub.github.io/SceneCraft
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。