arXiv:2505.19868cs.CV2025-05

将无训练技巧动态应用于3D生成,提升细节与表面质量

Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling

  • 在分数蒸馏采样中动态调整无训练技巧的强度
  • 平衡纹理细节与表面平滑度,减少几何错误
  • 适合追求高质量3D生成的开发者和研究者

近期研究表明,无训练技术(如无分类器引导、FreeU)能显著提升文本到2D生成质量。然而这些技术在分数蒸馏采样(SDS)中的作用尚未被充分探索,而SDS是利用预训练文本到2D扩散模型完成多种任务的常用方法。本文聚焦于通过2D升维实现文本到3D生成,研究无训练技术对SDS的影响。结果表明,调节无分类器引导(CFG)尺度在物体大小与表面平滑度间存在权衡;调节FreeU尺度则在纹理细节与几何误差间形成权衡。基于此,我们提出一种随时间步或优化迭代动态调整技巧强度的策略,有效平衡纹理细节与表面平滑度,保持输出尺寸并减少几何缺陷。

原文摘要 · Abstract (English)

Recent studies show that simple training-free techniques can dramatically improve the quality of text-to-2D generation outputs, e.g. Classifier-Free Guidance (CFG) or FreeU. However, these training-free techniques have been underexplored in the lens of Score Distillation Sampling (SDS), which is a popular and effective technique to leverage the power of pretrained text-to-2D diffusion models for various tasks. In this paper, we aim to shed light on the effect such training-free techniques have on SDS, via a particular application of text-to-3D generation via 2D lifting. We present our findings, which show that varying the scales of CFG presents a trade-off between object size and surface smoothness, while varying the scales of FreeU presents a trade-off between texture details and geometric errors. Based on these findings, we provide insights into how we can effectively harness training-free techniques for SDS, via a strategic scaling of such techniques in a dynamic manner with respect to the timestep or optimization iteration step. We show that using our proposed scheme strikes a favorable balance between texture details and surface smoothness in text-to-3D generations, while preserving the size of the output and mitigating the occurrence of geometric defects.

3D生成扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。