用分拆扩散模型实现可控制的3D场景生成与编辑
SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
- 将语义与几何信息分拆建模,通过语义盒控制场景生成
- 支持文本驱动生成,且可精准修改物体位置大小
- 适合需要精细调控3D场景的设计师或内容创作者
我们提出SceneFactor,一种基于扩散模型的大规模3D场景生成方法,支持可控生成与便捷编辑。该方法通过分拆扩散机制,利用潜在语义与几何流形生成任意尺寸的3D场景。文本输入可实现简便、可控的生成,但对局部编辑仍不够精确。为此,我们设计了分拆语义扩散,构建由语义3D盒组成的代理语义空间,通过增删或调整这些盒体,实现高保真、一致的3D几何编辑。大量实验表明,该方法在保持高质量生成的同时,实现了有效的可控编辑。
原文摘要 · Abstract (English)
We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion formulation, leveraging latent semantic and geometric manifolds for generation of arbitrary-sized 3D scenes. While text input enables easy, controllable generation, text guidance remains imprecise for intuitive, localized editing and manipulation of the generated 3D scenes. Our factored semantic diffusion generates a proxy semantic space composed of semantic 3D boxes that enables controllable editing of generated scenes by adding, removing, changing the size of the semantic 3D proxy boxes that guides high-fidelity, consistent 3D geometric editing. Extensive experiments demonstrate that our approach enables high-fidelity 3D scene synthesis with effective controllable editing through our factored diffusion approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。