arXiv:2412.20473cs.CV2024-12被引 3

用场景图与布局引导,生成复杂3D场景更精准。

Toward Scene Graph and Layout Guided Complex 3D Scene Generation

论文配图:Toward Scene Graph and Layout Guided Complex 3D Scene Generation
图 1 · 摘自论文原文
  • 通过场景图和超节点建模物体关系与交互
  • 生成结果更贴合文本描述,减少物体外观混淆
  • 适合需要精确控制多物体布局的3D内容创作

近期基于对象中心的文本到3D生成取得显著进展,但复杂3D场景生成仍面临挑战,主要源于物体间复杂关系。现有方法多依赖得分蒸馏采样(SDS),难以实现对多个物体及其特定交互的精细操控。为此,我们提出一种新框架——场景图与布局引导的3D场景生成(GraLa3D)。给定描述复杂3D场景的文本提示,GraLa3D利用大语言模型(LLM)构建包含布局边界框信息的场景图。该框架创新性地使用单物体节点与复合超节点共同构成场景图,并在超节点中建模物体间的交互关系,同时缓解同一节点内物体之间的外观泄漏问题。实验表明,GraLa3D有效克服了上述局限,生成的3D场景与文本提示高度一致。

原文摘要 · Abstract (English)

Recent advancements in object-centric text-to-3D generation have shown impressive results. However, generating complex 3D scenes remains an open challenge due to the intricate relations between objects. Moreover, existing methods are largely based on score distillation sampling (SDS), which constrains the ability to manipulate multiobjects with specific interactions. Addressing these critical yet underexplored issues, we present a novel framework of Scene Graph and Layout Guided 3D Scene Generation (GraLa3D). Given a text prompt describing a complex 3D scene, GraLa3D utilizes LLM to model the scene using a scene graph representation with layout bounding box information. GraLa3D uniquely constructs the scene graph with single-object nodes and composite super-nodes. In addition to constraining 3D generation within the desirable layout, a major contribution lies in the modeling of interactions between objects in a super-node, while alleviating appearance leakage across objects within such nodes. Our experiments confirm that GraLa3D overcomes the above limitations and generates complex 3D scenes closely aligned with text prompts.

3D生成场景图布局引导多物体交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。