通过渐进式几何优化与领域得分蒸馏,提升文本生成3D模型的细节与真实感。
DreamPolish: Domain Score Distillation With Progressive Geometry Generation
- 多神经表示结合视角条件扩散,用法向估计器逐步打磨几何表面。
- 仅需少量训练步骤即可显著减少几何瑕疵,生成更精细的3D结构。
- 引入领域得分蒸馏技术,使纹理更逼真且在预训练图像模型中保持一致。
我们提出DreamPolish,一种文本到3D生成模型,擅长生成精细几何与高质量纹理。在几何构建阶段,利用多种神经表示增强合成稳定性,不依赖单一视图条件扩散先验,而是引入额外的法向估计器,在不同视场角下优化几何细节。提出一个仅需少量训练步骤的表面精炼阶段,有效修复前序阶段引导不足导致的伪影,生成更具可实现性的3D物体。在纹理生成阶段,关键挑战在于从预训练文本到图像模型的巨大潜在分布中找到包含逼真且一致渲染结果的合适领域。为此,我们提出一种新的得分蒸馏目标——领域得分蒸馏(DSD),受无分类器指导(CFG)启发,表明CFG与变分分布引导代表梯度引导中不同的重要方面,二者共同提升纹理质量。大量实验表明,该模型生成的3D资产具有光滑表面和逼真纹理,优于现有最先进方法。
原文摘要 · Abstract (English)
We introduce DreamPolish, a text-to-3D generation model that excels in producing refined geometry and high-quality textures. In the geometry construction phase, our approach leverages multiple neural representations to enhance the stability of the synthesis process. Instead of relying solely on a view-conditioned diffusion prior in the novel sampled views, which often leads to undesired artifacts in the geometric surface, we incorporate an additional normal estimator to polish the geometry details, conditioned on viewpoints with varying field-of-views. We propose to add a surface polishing stage with only a few training steps, which can effectively refine the artifacts attributed to limited guidance from previous stages and produce 3D objects with more desirable geometry. The key topic of texture generation using pretrained text-to-image models is to find a suitable domain in the vast latent distribution of these models that contains photorealistic and consistent renderings. In the texture generation phase, we introduce a novel score distillation objective, namely domain score distillation (DSD), to guide neural representations toward such a domain. We draw inspiration from the classifier-free guidance (CFG) in textconditioned image generation tasks and show that CFG and variational distribution guidance represent distinct aspects in gradient guidance and are both imperative domains for the enhancement of texture quality. Extensive experiments show our proposed model can produce 3D assets with polished surfaces and photorealistic textures, outperforming existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。