arXiv:2409.11406cs.CV2024-09被引 26

用3D参考模型引导生成,提升细节质量与可控性。

Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion

  • 通过动态调节条件强度,融合图像与3D参考引导生成
  • 支持图像+3D参考输入,生成结果更贴近真实结构
  • 适合需要高精度3D建模的设计师或内容创作者

在3D建模中,设计师常以现有3D模型为参考来创建新模型。受此启发,我们提出Phidias,一种基于扩散模型的参考增强型3D生成方法。给定一张图像,该方法利用检索到或用户提供的3D参考模型来指导生成过程,从而提升生成质量、泛化能力和可控性。模型包含三个核心组件:1)元-ControlNet 动态调节条件强度;2)动态参考路由缓解输入图像与3D参考之间的错位问题;3)自参考增强技术实现渐进式自监督训练。这些设计共同推动了生成性能的显著提升。Phidias建立了一个统一框架,可同时处理文本、图像和3D条件,具备多样化应用潜力。

原文摘要 · Abstract (English)

In 3D modeling, designers often use an existing 3D model as a reference to create new ones. This practice has inspired the development of Phidias, a novel generative model that uses diffusion for reference-augmented 3D generation. Given an image, our method leverages a retrieved or user-provided 3D reference model to guide the generation process, thereby enhancing the generation quality, generalization ability, and controllability. Our model integrates three key components: 1) meta-ControlNet that dynamically modulates the conditioning strength, 2) dynamic reference routing that mitigates misalignment between the input image and 3D reference, and 3) self-reference augmentations that enable self-supervised training with a progressive curriculum. Collectively, these designs result in a clear improvement over existing methods. Phidias establishes a unified framework for 3D generation using text, image, and 3D conditions with versatile applications.

3D生成扩散模型参考引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。