用新损失函数提升文本生成3D的多样性与质量
Dive3D: Diverse Distillation-based Text-to-3D Generation via Score Implicit Matching
- 改用分数隐式匹配损失,避免模式崩溃
- 生成结果在多样性和视觉真实感上显著提升
- 适合追求高多样性3D生成的研究与应用
将预训练2D扩散模型蒸馏为3D资产已推动文本到3D生成的显著进展。然而,现有方法通常依赖基于KL散度的分数蒸馏采样(SDS)损失,其不对称结构天然偏向模式搜索,限制生成多样性。本文提出Dive3D,一种新框架,以分数隐式匹配(SIM)损失替代基于KL的目标,有效缓解模式崩溃。Dive3D在统一分歧视角下融合扩散蒸馏与奖励引导优化。该重构结合SIM损失,在保持文本对齐的同时,显著提升生成3D资产的多样性、人类偏好和整体视觉保真度。我们在多种2D到3D提示下验证Dive3D,定性评估显示其优于先前方法,涵盖多样性、照片真实感与美学吸引力。进一步在GPTEval3D基准测试中,对比九个先进基线,Dive3D在文本-资产对齐、3D合理性、文本-几何一致性、纹理质量与几何细节等量化指标上均表现优异。
原文摘要 · Abstract (English)
Distilling pre-trained 2D diffusion models into 3D assets has driven remarkable advances in text-to-3D synthesis. However, existing methods typically rely on Score Distillation Sampling (SDS) loss, which involves asymmetric KL divergence--a formulation that inherently favors mode-seeking behavior and limits generation diversity. In this paper, we introduce Dive3D, a novel text-to-3D generation framework that replaces KL-based objectives with Score Implicit Matching (SIM) loss, a score-based objective that effectively mitigates mode collapse. Furthermore, Dive3D integrates both diffusion distillation and reward-guided optimization under a unified divergence perspective. Such reformulation, together with SIM loss, yields significantly more diverse 3D outputs while improving text alignment, human preference, and overall visual fidelity. We validate Dive3D across various 2D-to-3D prompts and find that it consistently outperforms prior methods in qualitative assessments, including diversity, photorealism, and aesthetic appeal. We further evaluate its performance on the GPTEval3D benchmark, comparing against nine state-of-the-art baselines. Dive3D also achieves strong results on quantitative metrics, including text-asset alignment, 3D plausibility, text-geometry consistency, texture quality, and geometric detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。