arXiv:2502.05874cs.CVcs.AI2025-02AAAI被引 29

用图文混合图结构实现可精准控制几何的3D室内场景生成

MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation

  • 构建图文混合图结构,支持文本与视觉模态融合输入
  • 通过关系预测补全缺失节点关系,提升场景布局一致性
  • 适合需要精细几何控制的虚拟现实与室内设计场景

可控3D场景生成在虚拟现实和室内设计中有广泛应用,要求生成场景在几何上具备高真实感与可控性。场景图提供合适的数据表示形式。然而,现有基于图的方法仅支持文本输入,对灵活用户输入适应性差,难以精确控制物体几何。为此,我们提出MMGDreamer,一种双分支扩散模型,包含新颖的混合模态图、视觉增强模块和关系预测器。混合模态图使物体节点能融合文本与视觉模态,并支持可选节点间关系,提升对用户输入的适应性,实现对生成场景中物体几何的精细控制。视觉增强模块利用文本嵌入构建文本节点的视觉表示,提升视觉保真度。关系预测器基于节点表示推断缺失关系,生成更连贯的场景布局。大量实验表明,MMGDreamer在物体几何控制方面表现优异,达到当前最优场景生成性能。

原文摘要 · Abstract (English)

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. However, current graph-based methods for scene generation are constrained to text-based inputs and exhibit insufficient adaptability to flexible user inputs, hindering the ability to precisely control object geometry. To address this issue, we propose MMGDreamer, a dual-branch diffusion model for scene generation that incorporates a novel Mixed-Modality Graph, visual enhancement module, and relation predictor. The mixed-modality graph allows object nodes to integrate textual and visual modalities, with optional relationships between nodes. It enhances adaptability to flexible user inputs and enables meticulous control over the geometry of objects in the generated scenes. The visual enhancement module enriches the visual fidelity of text-only nodes by constructing visual representations using text embeddings. Furthermore, our relation predictor leverages node representations to infer absent relationships between nodes, resulting in more coherent scene layouts. Extensive experimental results demonstrate that MMGDreamer exhibits superior control of object geometry, achieving state-of-the-art scene generation performance. Project page: https://yangzhifeio.github.io/project/MMGDreamer.

3D生成图文融合场景控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。