arXiv:2512.04985stat.MLcs.LG2025-12被引 6

统一了扩散模型的引导生成框架,揭示了分类器无关引导的优化目标。

Towards a unified framework for guided diffusion models

  • 提出将奖励梯度注入反向扩散过程,统一处理各类引导机制。
  • 理论证明分类器无关引导能降低分类器概率的期望倒数。
  • 新采样器训练简单,无需完整扩散轨迹,适合实际应用。

带引导或控制的数据生成已成为现代生成建模的核心。尽管扩散模型理论取得显著进展,但引导扩散采样器的理论理解仍严重不足。本文构建了一个统一的算法与理论框架,涵盖扩散引导和奖励引导两种情形。针对微调扩散模型以提升特定奖励的目标,我们提出在反向扩散过程中注入一个由原始与奖励重加权得分差构成的奖励引导项,并严格量化其相对于无引导情况的奖励提升效果。作为关键应用,该框架首次理论刻画了分类器无关引导(CFG)所优化的具体性能指标:对一般目标分布,它能降低分类器概率的期望倒数。应用于奖励引导扩散时,框架提出一种易于训练的新采样器,训练阶段无需完整扩散轨迹。数值实验进一步验证了理论结果。

原文摘要 · Abstract (English)

Guided or controlled data generation with diffusion models\blfootnote{Partial preliminary results of this work appeared in International Conference on Machine Learning 2025 \citep{li2025provable}.} has become a cornerstone of modern generative modeling. Despite substantial advances in diffusion model theory, the theoretical understanding of guided diffusion samplers remains severely limited. We make progress by developing a unified algorithmic and theoretical framework that accommodates both diffusion guidance and reward-guided diffusion. Aimed at fine-tuning diffusion models to improve certain rewards, we propose injecting a reward guidance term -- constructed from the difference between the original and reward-reweighted scores -- into the backward diffusion process, and rigorously quantify the resulting reward improvement over the unguided counterpart. As a key application, our framework shows that classifier-free guidance (CFG) decreases the expected reciprocal of the classifier probability, providing the first theoretical characterization of the specific performance metric that CFG improves for general target distributions. When applied to reward-guided diffusion, our framework yields a new sampler that is easy-to-train and requires no full diffusion trajectories during training. Numerical experiments further corroborate our theoretical findings.

扩散模型引导生成理论分析奖励引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。