用3D布局引导生成更真实的3D场景,精准控制物体位置。
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
- 通过语义布局分解场景,分阶段优化几何与外观。
- 在多个数据集上生成效果优于现有方法,更符合真实场景。
- 适合需要精确场景编辑的应用,如游戏、影视制作。
3D场景生成得益于2D扩散模型的发展已取得显著进展。然而,文本对3D场景的描述本身不准确且训练中缺乏细粒度控制,导致生成结果不合理。为解决此问题,我们提出一种基于3D语义布局提示的文本到场景生成方法(Layout2Scene),通过额外布局信息精确控制物体位置。首先引入混合场景表示,分离物体与背景,并利用预训练的文本到3D模型初始化。随后采用两阶段方案分别优化几何与外观。为充分利用2D扩散先验,我们设计了语义引导的几何扩散模型和语义-几何联合引导的扩散模型,均在场景数据集上微调。大量实验表明,该方法生成的场景比现有先进方法更合理、更逼真,且支持灵活精确编辑,适用于多种下游应用。
原文摘要 · Abstract (English)
3D scene generation conditioned on text prompts has significantly progressed due to the development of 2D diffusion generation models. However, the textual description of 3D scenes is inherently inaccurate and lacks fine-grained control during training, leading to implausible scene generation. As an intuitive and feasible solution, the 3D layout allows for precise specification of object locations within the scene. To this end, we present a text-to-scene generation method (namely, Layout2Scene) using additional semantic layout as the prompt to inject precise control of 3D object positions. Specifically, we first introduce a scene hybrid representation to decouple objects and backgrounds, which is initialized via a pre-trained text-to-3D model. Then, we propose a two-stage scheme to optimize the geometry and appearance of the initialized scene separately. To fully leverage 2D diffusion priors in geometry and appearance generation, we introduce a semantic-guided geometry diffusion model and a semantic-geometry guided diffusion model which are finetuned on a scene dataset. Extensive experiments demonstrate that our method can generate more plausible and realistic scenes as compared to state-of-the-art approaches. Furthermore, the generated scene allows for flexible yet precise editing, thereby facilitating multiple downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。