用2D扩散模型直接生成3D高保真物体,无需3D数据训练。
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
- 将3D高斯点云拆解为多视角图像,借力2D扩散模型生成。
- 引入跨视角与跨属性注意力,确保生成点云的3D一致性。
- 首次实现无需3D数据训练的端到端图像到3D生成,适合新物体泛化。
2D图像生成近年来取得显著进展,主要得益于扩散模型的强大能力与大规模数据集的支持。然而,直接进行3D生成仍受限于3D数据稀缺且质量较低。本文提出Zero-1-to-G,一种利用预训练2D扩散模型直接生成高斯点云的新方法。核心思路是将3D高斯点云分解为编码不同属性的多视角图像,从而将复杂的3D生成任务重构为2D扩散框架中的问题,充分借助预训练2D扩散模型的丰富先验。为引入3D感知,我们设计了跨视角与跨属性注意力层,捕捉复杂相关性并强制生成点云的3D一致性。该方法是首个有效利用预训练2D扩散先验的直接图像到3D生成模型,支持高效训练并提升对未见物体的泛化能力。在合成与真实场景数据集上的大量实验表明,其在3D物体生成方面性能卓越,为高质量3D生成提供了新路径。
原文摘要 · Abstract (English)
Recent advances in 2D image generation have achieved remarkable quality,largely driven by the capacity of diffusion models and the availability of large-scale datasets. However, direct 3D generation is still constrained by the scarcity and lower fidelity of 3D datasets. In this paper, we introduce Zero-1-to-G, a novel approach that addresses this problem by enabling direct single-view generation on Gaussian splats using pretrained 2D diffusion models. Our key insight is that Gaussian splats, a 3D representation, can be decomposed into multi-view images encoding different attributes. This reframes the challenging task of direct 3D generation within a 2D diffusion framework, allowing us to leverage the rich priors of pretrained 2D diffusion models. To incorporate 3D awareness, we introduce cross-view and cross-attribute attention layers, which capture complex correlations and enforce 3D consistency across generated splats. This makes Zero-1-to-G the first direct image-to-3D generative model to effectively utilize pretrained 2D diffusion priors, enabling efficient training and improved generalization to unseen objects. Extensive experiments on both synthetic and in-the-wild datasets demonstrate superior performance in 3D object generation, offering a new approach to high-quality 3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。