用多视角扩散模型生成可控的3D室内场景,支持文本到场景生成。
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
- 基于粗略3D布局引导多视角图像生成,保持视角一致性。
- 引入布局感知极线注意力机制,提升生成场景的多视图一致性。
- 支持迭代生成不同物体数量与复杂度的场景,适合场景设计应用。
我们提出MVRoom,一种可控制的新视角合成(NVS)管道,用于生成3D室内场景,采用多视角扩散模型并以粗略3D布局为条件。该方法采用两阶段设计:第一阶段通过新型表示,有效连接3D布局与一致的图像条件信号,实现多视角生成;第二阶段在图像条件基础上进行多视角生成,并引入布局感知极线注意力机制,在扩散过程中增强多视图一致性。此外,我们设计了一种迭代框架,通过递归执行多视角生成(MVRoom),支持生成包含不同物体数量与场景复杂度的3D场景,实现文本到场景生成。实验结果表明,该方法在定量与定性评估上均优于现有最先进基线方法。消融实验进一步验证了生成流程中关键组件的有效性。
原文摘要 · Abstract (English)
We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-stage design in which the 3D layout is used throughout to enforce multi-view consistency. The first stage employs novel representations to effectively bridge the 3D layout and consistent image-based condition signals for multi-view generation. The second stage performs image-conditioned multi-view generation, incorporating a layout-aware epipolar attention mechanism to enhance multi-view consistency during the diffusion process. Additionally, we introduce an iterative framework that generates 3D scenes with varying numbers of objects and scene complexities by recursively performing multi-view generation (MVRoom), supporting text-to-scene generation. Experimental results demonstrate that our approach achieves high-fidelity and controllable 3D scene generation for NVS, outperforming state-of-the-art baseline methods both quantitatively and qualitatively. Ablation studies further validate the effectiveness of key components within our generation pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。