arXiv:2506.13763cs.LGcs.AI2025-06被引 5

通过估算最优损失值,提升扩散模型训练诊断与优化效果。

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value

  • 提出统一框架下计算扩散模型最优损失的闭式解及可扩展估计方法。
  • 发现减去最优损失后,120M至1.5B参数模型更符合幂律规律。
  • 为训练质量评估和优化提供新指标,适用于主流扩散模型研究。

扩散模型在生成建模中取得显著成功,但其训练损失无法直接反映数据拟合质量,因为最优损失通常非零且未知,导致大最优损失与模型容量不足难以区分。本文提出需估算最优损失以诊断和改进扩散模型。我们基于统一框架推导出最优损失的闭式解,并开发了有效估计器,包括可扩展至大规模数据集、可控方差与偏差的随机变体。借助该工具,我们建立了主流扩散模型变体的内在训练质量评估指标,并设计了更优的训练调度策略。进一步地,使用120M至1.5B参数模型发现,将实际训练损失减去最优损失后,幂律关系更明显,为扩散模型的缩放规律研究提供了更合理的设定。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not indicative of absolute data-fitting quality, since its optimal value is typically not zero but unknown, leading to confusion between large optimal loss and insufficient model capacity. In this work, we advocate the need to estimate the optimal loss value for diagnosing and improving diffusion models. We first derive the optimal loss in closed form under a unified formulation of diffusion models, and develop effective estimators for it, including a stochastic variant scalable to large datasets with proper control of variance and bias. With this tool, we unlock the inherent metric for diagnosing the training quality of mainstream diffusion model variants, and develop a more performant training schedule based on the optimal loss. Moreover, using models with 120M to 1.5B parameters, we find that the power law is better demonstrated after subtracting the optimal loss from the actual training loss, suggesting a more principled setting for investigating the scaling law for diffusion models.

扩散模型损失估计训练诊断缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。