arXiv:2607.02407cs.AIcs.CV2026-07

让AI在非曼哈顿空间生成更真实的室内场景

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

论文配图:Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
图 1 · 摘自论文原文
  • 用统计先验和分层布局提升复杂空间理解
  • 新基准测试显示显著优于现有方法
  • 适合做真实场景重建与虚拟设计的团队

大型语言模型在曼哈顿环境下的3D室内生成已取得显著进展,但现有方法难以捕捉非曼哈顿场景中合理的物体布局,主要因无法建模非正交空间关系,导致几何错误多、物理真实性低。为此,我们提出SPG-Layout,一种面向文本驱动的非曼哈顿环境3D室内场景生成框架。首先利用物体分布的统计先验指导训练,增强环境理解与保真度;其次借鉴人类设计流程,采用分层布局策略,优先放置大件物体,大幅减少布局违规。二者协同实现语义真实与物理合理性的平衡。为评估复杂场景表现,我们构建了包含500个多样化非曼哈顿环境的新基准。大量实验表明,SPG-Layout在曼哈顿与非曼哈顿环境中均显著优于现有方法。代码将公开发布。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statistical priors of object distributions to guide the training process, enhancing environmental understanding and fidelity. Furthermore, mirroring human design workflows, we adopt a hierarchical layout strategy that prioritizes the placement of large objects, thereby substantially minimizing layout violations. By synergizing these components, SPG-Layout achieves a balanced optimization of semantic realism and physical plausibility. To evaluate performance in these complex settings, we constructed a new benchmark comprising 500 diverse non-Manhattan environments. Extensive experiments demonstrate that SPG-Layout consistently and significantly outperforms existing methods across both Manhattan and non-Manhattan environments. The code will be publicly released.

3D生成文本生成布局优化非曼哈顿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。