分阶段生成3D室内布局,让家具摆放更合理且可控。
CasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling

- 分四步逐步生成:先定家具数量类别,再调大小与特征,建空间关系,最后出边界框。
- 在复杂平面中保持物理合理性,生成布局更符合功能组织,比现有方法更逼真多样。
- 用稀疏关系图模拟人类对空间的描述习惯,适合需要零样本控制的场景生成任务。
由于数据稀缺以及难以同时满足全局建筑约束和局部语义一致性,合成真实3D室内场景仍具挑战。现有方法常忽略结构边界或依赖全连接关系图,导致冗余生成误差。受人类设计认知启发,我们提出CasLayout,一种级联扩散框架,将联合场景生成任务分解为四个条件子阶段:(1) 预测家具数量与类别,(2) 优化物体尺寸与特征嵌入,(3) 在隐空间建模空间关系,(4) 生成定向边界框(OBB)。该解耦架构降低数据需求,并支持灵活集成大语言模型(LLMs)与视觉语言模型(VLMs),实现图像到场景等零样本任务。为确保复杂平面内的物理有效性,显式将墙体、门、窗等建筑元素作为条件约束。针对密集关系图的高熵问题,引入与人类空间描述一致的稀疏关系图形式。通过双向变分自编码器(VAE)将该稀疏图编码至紧凑隐空间,提升关系可控性,使生成布局更符合功能组织。实验表明,CasLayout在保真度与多样性上达到当前最优水平,且在实际应用中具备更强可控性。
原文摘要 · Abstract (English)
Synthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural constraints and local semantic consistency. Existing approaches often overlook structural boundaries or rely on fully connected relation graphs that introduce redundant generation errors. Inspired by human design cognition, we present CasLayout, a cascaded diffusion framework that decomposes the joint scene generation task into four conditional sub-stages with explicit physical and semantic roles: (1) predicting furniture quantity and categories, (2) refining object sizes and feature embeddings, (3) modeling spatial relationships in a latent space, and (4) generating Oriented Bounding Boxes (OBBs). This decoupled architecture reduces data requirements and enables flexible integration of Large Language Models (LLMs) and Vision Language Models (VLMs) for zero-shot tasks such as image-to-scene generation. To maintain physical validity within complex floor plans, we explicitly model building elements (e.g., walls, doors, and windows) as conditional constraints. Furthermore, to address the high entropy of dense relation graphs, we introduce a sparse relation graph formulation aligned with human spatial descriptions. By encoding these sparse graphs into a compact latent space using a bidirectional Variational Autoencoder (VAE), the proposed framework provides enhanced relational controllability, allowing generated layouts to better respect functional organization. Experiments demonstrate that CasLayout achieves state-of-the-art performance in fidelity and diversity while enabling improved controllability in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。