用物理规则指导文本生成3D场景布局,让多个物体摆放更真实可控。
LAYOUTDREAMER: Physics-guided Layout for Text-to-3D Compositional Scene Generation

- 通过文本转为有向场景图,动态调整3D高斯点密度与位置
- 在T3Bench上多物体生成指标达当前最优(SOTA)
- 适合需要真实物理布局的3D内容生成任务
近期,文本引导的3D场景生成领域受到广泛关注。高质量且符合物理真实的生成结果对实际应用至关重要。然而现有方法存在三大局限:(i) 难以捕捉文本描述中多个物体间的复杂关系,(ii) 无法生成物理上合理的场景布局,(iii) 在组合场景中缺乏可控性与可扩展性。本文提出LayoutDreamer,利用3D高斯溅射(3DGS)实现高质量、物理一致的组合场景生成。给定文本提示后,先将其转化为有向场景图,并自适应调整初始组合3D高斯的密度与布局;随后基于训练焦距动态调整相机,确保实体级生成质量;最后,从场景图中提取有向依赖关系,定制物理与布局能量项,兼顾真实性与灵活性。大量实验表明,LayoutDreamer在组合场景生成质量与语义对齐方面均优于现有方法,在T3Bench的多物体生成指标上达到最先进(SOTA)水平。
原文摘要 · Abstract (English)
Recently, the field of text-guided 3D scene generation has garnered significant attention. High-quality generation that aligns with physical realism and high controllability is crucial for practical 3D scene applications. However, existing methods face fundamental limitations: (i) difficulty capturing complex relationships between multiple objects described in the text, (ii) inability to generate physically plausible scene layouts, and (iii) lack of controllability and extensibility in compositional scenes. In this paper, we introduce LayoutDreamer, a framework that leverages 3D Gaussian Splatting (3DGS) to facilitate high-quality, physically consistent compositional scene generation guided by text. Specifically, given a text prompt, we convert it into a directed scene graph and adaptively adjust the density and layout of the initial compositional 3D Gaussians. Subsequently, dynamic camera adjustments are made based on the training focal point to ensure entity-level generation quality. Finally, by extracting directed dependencies from the scene graph, we tailor physical and layout energy to ensure both realism and flexibility. Comprehensive experiments demonstrate that LayoutDreamer outperforms other compositional scene generation quality and semantic alignment methods. Specifically, it achieves state-of-the-art (SOTA) performance in the multiple objects generation metric of T3Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。