通过一致性流蒸馏,提升文本到3D生成的视觉质量与多样性。
Consistent Flow Distillation for Text-to-3D Generation
- 基于扩散过程梯度,引入多视角一致高斯噪声引导3D生成。
- 在多个视角下保持图像流一致性,显著提升3D生成质量。
- 适合追求高质量3D内容生成的研究者与开发者。
Score Distillation Sampling (SDS) 在将图像生成模型蒸馏用于3D生成方面取得了显著进展。然而,其最大似然导向的行为常导致视觉质量下降和多样性不足,限制了在3D应用中的有效性。本文提出一致流蒸馏(CFD),通过利用扩散ODE或SDE采样过程的梯度来指导3D生成。从梯度采样视角出发,我们发现不同视角间2D图像流的一致性对高质量3D生成至关重要。为此,我们在3D物体上引入多视角一致的高斯噪声,可从多个视角渲染并计算流梯度。实验表明,通过一致流机制,CFD在文本到3D生成任务中显著优于以往方法。
原文摘要 · Abstract (English)
Score Distillation Sampling (SDS) has made significant strides in distilling image-generative models for 3D generation. However, its maximum-likelihood-seeking behavior often leads to degraded visual quality and diversity, limiting its effectiveness in 3D applications. In this work, we propose Consistent Flow Distillation (CFD), which addresses these limitations. We begin by leveraging the gradient of the diffusion ODE or SDE sampling process to guide the 3D generation. From the gradient-based sampling perspective, we find that the consistency of 2D image flows across different viewpoints is important for high-quality 3D generation. To achieve this, we introduce multi-view consistent Gaussian noise on the 3D object, which can be rendered from various viewpoints to compute the flow gradient. Our experiments demonstrate that CFD, through consistent flows, significantly outperforms previous methods in text-to-3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。