区分主次物体,分两阶段生成更真实的密集室内场景。
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation

- 按物体角色分主次,拆解为结构与上下文两阶段生成。
- 主物体先生成全局一致的布局骨架,提升物理合理性。
- 适合需要高保真仿真环境的具身AI研究者使用。
生成可控制且符合物理规律的室内场景是构建高保真模拟环境以支持具身AI的关键前提。然而,现有基于深度学习的方法通常将所有物体视为统一生成过程中的同质实例。尽管在稀疏简单布局上表现良好,但在密集物体排列和复杂空间依赖关系下难以建模,导致可扩展性差且物理合理性下降。为此,我们从结构异质性的角度重新审视室内布局生成,根据物体在场景塑造中的不同作用,将物体分解为主物体与次物体。基于此,提出HetScene:一种异质性两阶段生成框架,将室内布局合成解耦为结构布局生成(SLG)与上下文布局生成(CLG)。SLG首先仅基于文本描述、自上而下的二值房间掩码及空间关系图,生成仅含主物体的全局一致结构布局,建立大型核心家具的稳定全局宏观骨架。
原文摘要 · Abstract (English)
Generating controllable and physically plausible indoor scenes is a pivotal prerequisite for constructing high-fidelity simulation environments for embodied AI. However, existing deeplearning-based methods usually treat all objects as homogeneous instances within a unified generation process. While effective for sparse and simplistic layouts, they struggle to model realistic layouts with dense object arrangements and complex spatial dependencies, leadingto limited scalability and degraded physical plausibility. To deal with these challenges, we revisit indoor layout generation from the perspective of structural heterogeneity and decompose the objects into primary objects and secondary objects according to their distinct roles in shaping a scene. Based on this decomposition, we propose HetScene, a heterogeneous two-stage generation framework that decouples indoor layout synthesis into Structural Layout Generation (SLG) and Contextual Layout Generation (CLG). SLG first generates globally coherent structural layouts with only primary objects conditioned on text descriptions, top-down binary room masks, and spatial relation graphs, establishing a stable global macro-skeleton of large core furniture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。