arXiv:2502.00708cs.CVcs.AI2025-02被引 7

用物理约束生成更合理的3D组合场景,效率提升24倍。

PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation

  • 结合大语言模型与世界模型,分步生成场景图与3D资产。
  • 在CLIP分数上达到当前最佳,生成质量媲美顶尖方法。
  • 适合需要高物理合理性与高效生成的3D内容创作者。

文本生成3D资产虽借助2D扩散先验取得进展,但在组合场景生成中仍面临三大挑战:难以保证布局符合物理规律;难以准确捕捉复杂描述中的物体及其关系;基于大语言模型的布局方法自主生成能力有限。为此,我们提出新框架PhiP-G,通过融合生成技术与基于世界模型的布局引导实现突破。利用大语言模型代理分析复杂场景描述生成场景图,结合多模态2D生成代理与3D高斯生成方法进行精准资产创建。布局阶段则引入具备附着特性的物理池与视觉监督代理,构建世界模型以预测和规划布局。大量实验表明,PhiP-G显著提升组合场景的生成质量与物理合理性。其在CLIP得分上达到当前最优,在T$^3$Bench上生成质量与领先方法相当,且效率提升24倍。

原文摘要 · Abstract (English)

Text-to-3D asset generation has achieved significant optimization under the supervision of 2D diffusion priors. However, when dealing with compositional scenes, existing methods encounter several challenges: 1). failure to ensure that composite scene layouts comply with physical laws; 2). difficulty in accurately capturing the assets and relationships described in complex scene descriptions; 3). limited autonomous asset generation capabilities among layout approaches leveraging large language models (LLMs). To avoid these compromises, we propose a novel framework for compositional scene generation, PhiP-G, which seamlessly integrates generation techniques with layout guidance based on a world model. Leveraging LLM-based agents, PhiP-G analyzes the complex scene description to generate a scene graph, and integrating a multimodal 2D generation agent and a 3D Gaussian generation method for targeted assets creation. For the stage of layout, PhiP-G employs a physical pool with adhesion capabilities and a visual supervision agent, forming a world model for layout prediction and planning. Extensive experiments demonstrate that PhiP-G significantly enhances the generation quality and physical rationality of the compositional scenes. Notably, PhiP-G attains state-of-the-art (SOTA) performance in CLIP scores, achieves parity with the leading methods in generation quality as measured by the T$^3$Bench, and improves efficiency by 24x.

3D生成物理约束场景理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。