arXiv:2409.03718cs.CVcs.GR2024-09ICLR被引 18

用2D图像表示3D形状,实现快速文本生成3D模型

Geometry Image Diffusion: Fast and Data-Efficient Text-to-3D with Image-Based Surface Representation

  • 用几何图像替代复杂3D结构,以2D方式高效表示3D形状
  • 仅用高质量有限数据训练,生成速度媲美文生图模型
  • 支持分部件生成与内部结构,适合3D资产创作场景

从文本描述生成高质量3D物体仍面临计算成本高、3D数据稀缺及表示复杂等挑战。我们提出几何图像扩散(GIMDiffusion),一种新型文生3D模型,利用几何图像以2D图像形式高效表示3D形状,避免使用复杂的3D感知架构。通过引入协同控制机制,充分利用现有文生图模型(如Stable Diffusion)的丰富2D先验知识。该方法在仅使用高质量3D训练数据的情况下仍具备强泛化能力,并保持与IPAdapter等引导技术的兼容性。总体而言,GIMDiffusion可实现与当前文生图模型相当的生成速度,生成的3D对象包含语义明确的独立部件和内部结构,显著提升可用性与灵活性。

原文摘要 · Abstract (English)

Generating high-quality 3D objects from textual descriptions remains a challenging problem due to computational cost, the scarcity of 3D data, and complex 3D representations. We introduce Geometry Image Diffusion (GIMDiffusion), a novel Text-to-3D model that utilizes geometry images to efficiently represent 3D shapes using 2D images, thereby avoiding the need for complex 3D-aware architectures. By integrating a Collaborative Control mechanism, we exploit the rich 2D priors of existing Text-to-Image models such as Stable Diffusion. This enables strong generalization even with limited 3D training data (allowing us to use only high-quality training data) as well as retaining compatibility with guidance techniques such as IPAdapter. In short, GIMDiffusion enables the generation of 3D assets at speeds comparable to current Text-to-Image models. The generated objects consist of semantically meaningful, separate parts and include internal structures, enhancing both usability and versatility.

文生3D几何图像扩散模型快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。