arXiv:2502.06819cs.LGcs.GR2025-02被引 1

用图扩散生成3D场景,让物体摆放更合理、适合人使用。

AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics

  • 先用图扩散生成语义连贯的场景图,再布局物体。
  • 引入人物交互先验,减少物体重叠,提升实用性。
  • 适合需要真实家居布局的生成任务,如虚拟装修。

本文提出一种从文本提示生成3D室内场景的框架。现有方法通常将场景合成建模为仅依赖单一输入模态(如文本描述、房间形状或场景图)的对象布局预测问题,易导致物体碰撞且功能性不足,限制实际应用。为此,我们设计多阶段流程,更贴近真实场景创建过程:给定描述部分场景内容的文本提示,首先利用图扩散生成语义一致的场景图,随后预测合理的物体布局。此外,引入轻量级人-物交互先验,鼓励以人为核心的合理布置,并通过显式空间约束减少物体穿透。该方法生成的3D场景在布局上具有高一致性,更支持人类交互。在3D-FRONT数据集上的实验表明,本方法性能与现有方法相比达到竞争性或领先水平,同时显著提升了生成场景的物理合理性。

原文摘要 · Abstract (English)

This paper presents a framework for generating 3D indoor scenes from text prompts. Existing methods often formulate scene synthesis as an object layout prediction problem conditioned on a single input modality, such as a text description, room shape, or scene graph. This design can lead to object collisions and limited functional plausibility, reducing its practical applicability. To address these limitations, we introduce a multi-stage pipeline that better reflects practical scene creation scenarios. Given a text prompt describing partial scene content, our method first uses graph diffusion to produce a contextually coherent scene graph and then predicts a realistic object layout. In addition, we incorporate lightweight human-object interaction priors to encourage human-centric and functional arrangements, with explicit spatial constraints to reduce interpenetration. Our approach generates coherent 3D scenes with viable layouts that better support human interaction. Experiments on the 3D-FRONT dataset demonstrate that our method achieves competitive or state-of-the-art performance compared with existing approaches, while improving the physical plausibility of generated scenes.

3D生成场景合成图扩散人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。