arXiv:2509.14981cs.CV2025-09被引 16

用布局和参考图生成逼真且语义一致的3D室内场景。

SPATIALGEN: Layout-guided 3D Indoor Scene Generation

  • 基于多视角多模态扩散模型,结合布局与参考图生成3D场景。
  • 在12,328个结构化场景上训练,生成结果优于现有方法。
  • 适合做虚拟现实、建筑设计与机器人导航的研究者使用。

高质量3D室内环境建模对设计、虚拟现实和机器人应用至关重要,但人工建模耗时费力。尽管生成式AI已能实现自动化场景合成,现有方法仍难以平衡视觉质量、多样性、语义一致性与用户控制。主要瓶颈在于缺乏大规模高质量数据集。为此,我们构建了一个综合合成数据集,包含12,328个结构化标注场景、57,431个房间以及470万张照片级2D渲染图。基于该数据集,我们提出SpatialGen,一种新型多视角多模态扩散模型,可给定3D布局和参考图像(由文本提示生成),从任意视角合成外观(颜色图)、几何(场景坐标图)和语义(分割图),并保持跨模态空间一致性。实验表明,SpatialGen生成效果显著优于先前方法。我们开源数据与模型,以推动室内场景理解与生成领域发展。

原文摘要 · Abstract (English)

Creating high-fidelity 3D models of indoor environments is essential for applications in design, virtual reality, and robotics. However, manual 3D modeling remains time-consuming and labor-intensive. While recent advances in generative AI have enabled automated scene synthesis, existing methods often face challenges in balancing visual quality, diversity, semantic consistency, and user control. A major bottleneck is the lack of a large-scale, high-quality dataset tailored to this task. To address this gap, we introduce a comprehensive synthetic dataset, featuring 12,328 structured annotated scenes with 57,431 rooms, and 4.7M photorealistic 2D renderings. Leveraging this dataset, we present SpatialGen, a novel multi-view multi-modal diffusion model that generates realistic and semantically consistent 3D indoor scenes. Given a 3D layout and a reference image (derived from a text prompt), our model synthesizes appearance (color image), geometry (scene coordinate map), and semantic (semantic segmentation map) from arbitrary viewpoints, while preserving spatial consistency across modalities. SpatialGen consistently generates superior results to previous methods in our experiments. We are open-sourcing our data and models to empower the community and advance the field of indoor scene understanding and generation.

3D生成扩散模型室内场景多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。