arXiv:2603.22364cs.LGcs.AI2026-03被引 1

通过最大化类别间似然比,让扩散模型无需引导也能生成高质量条件样本。

MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives

  • 训练时显式最大化不同类别间的似然比,增强类别分离。
  • 无需推理时引导,生成质量接近传统引导方法,提升无引导生成效果。
  • 揭示了引导机制的本质是隐式对比对齐,为指导设计提供理论依据。

扩散模型在生成建模中表现优异,但其性能常依赖于推理时的分类器自由引导(CFG),这是一种修改采样轨迹的启发式方法。理论上,通过标准去噪得分匹配(DSM)训练的扩散模型应能恢复目标数据分布,这引出两个根本问题:(i) 为何实践中需要推理时引导?(ii) 其潜在效果能否内化为一个合理的训练目标?本文认为,标准DSM的一个关键局限是类别间分离不足。为此,我们提出MCLR,一种在训练中显式最大化类别间似然比的对齐目标。用MCLR微调扩散模型后,在标准采样下即可实现类似CFG的改进,显著提升无引导条件生成能力,并缩小与推理时引导之间的差距。此外,我们从理论上证明,带有CFG引导的得分恰好是样本自适应加权MCLR目标的最优解。这一结果将CFG与基于对齐的目标联系起来,揭示了CFG本质上是一种隐式的推理时对比对齐过程。

原文摘要 · Abstract (English)

Diffusion models achieve strong performance in generative modeling, but their success often relies heavily on classifier-free guidance (CFG), an inference-time heuristic that modifies the sampling trajectory. In theory, diffusion models trained with standard denoising score matching (DSM) should recover the target data distribution, raising two fundamental questions: (i) why is inference-time guidance necessary in practice, and (ii) can its underlying effect be internalized into a principled training objective? In this work, we argue that a key limitation of standard DSM is insufficient inter-class separation. To address this issue, we propose MCLR, an alignment objective that explicitly maximizes inter-class likelihood-ratios during training. Fine-tuning diffusion models with MCLR induces CFG-like improvements under standard sampling, substantially improving guidance-free conditional generation and narrowing the gap to inference-time CFG. Beyond these empirical benefits, we show theoretically that the CFG-guided score is exactly the optimal solution to a sample-adaptive weighted MCLR objective. This result connects CFG to alignment-based objectives, providing a mechanistic interpretation of CFG as an implicit inference-time contrastive alignment procedure.

扩散模型条件生成对齐学习无引导生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。