用2D扩散模型生成3D物体,仅需少量3D数据即可实现。
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
- 用密集2D网格构建高斯图谱,将2D模型迁移到3D空间。
- 在20.5万件3D物体上训练,成功实现文本到3D生成。
- 适合想低成本做3D生成的研究者与创作者。
近年来,文本到图像扩散模型的进步得益于大量配对的2D数据。然而,3D扩散模型因高质量3D数据稀缺而发展受限,性能远不及2D模型。为解决这一问题,我们提出复用预训练2D扩散模型进行3D物体生成。引入高斯图谱(Gaussian Atlas)这一新表示,利用密集2D网格,使2D扩散模型可微调生成3D高斯点。该方法实现了从预训练2D模型到3D结构扁平化2D流形的成功迁移学习。为支持训练,我们构建了包含20.5万件高质量3D高斯拟合物体的GaussianVerse数据集。实验表明,文本到图像扩散模型可有效适配于3D内容生成,显著缩小2D与3D建模之间的差距。
原文摘要 · Abstract (English)
Recent advances in text-to-image diffusion models have been driven by the increasing availability of paired 2D data. However, the development of 3D diffusion models has been hindered by the scarcity of high-quality 3D data, resulting in less competitive performance compared to their 2D counterparts. To address this challenge, we propose repurposing pre-trained 2D diffusion models for 3D object generation. We introduce Gaussian Atlas, a novel representation that utilizes dense 2D grids, enabling the fine-tuning of 2D diffusion models to generate 3D Gaussians. Our approach demonstrates successful transfer learning from a pre-trained 2D diffusion model to a 2D manifold flattened from 3D structures. To support model training, we compile GaussianVerse, a large-scale dataset comprising 205K high-quality 3D Gaussian fittings of various 3D objects. Our experimental results show that text-to-image diffusion models can be effectively adapted for 3D content generation, bridging the gap between 2D and 3D modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。