揭示条件流匹配何时可替代负对数似然,为模型训练提供理论依据。
When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?

- 通过分解终点负对数似然,揭示条件流匹配的误差来源。
- 仅在特定权重下(如 $w_{\mathrm{sc}}(t)=(1-t)/t$)可消除内部残差,实现精确估计。
- 适用于需控制近似误差的生成模型训练,尤其适合对比学习与对齐任务。
流匹配实现无似然训练,但现有方法常将条件流匹配(CFM)损失用作终点负对数似然(NLL),并以新旧差异作为似然比。本文刻画了这些替换的适用条件。对于线性高斯路径,我们精确分解终点NLL为熵、加权CFM目标、内部速度-得分残差和边界残差。因此,仅当对应残差抵消时,纯CFM估计与差异才是精确的。在非策略总体最优下,普通CFM通常不是点态NLL估计器;而权重 $w_{\mathrm{sc}}(t)=(1-t)/t$ 可消除内部残差;该正结果不普遍适用于训练或策略内对齐。即使终点分布相同或经代理优化,策略内似然比仍可能有偏。跨维度、分布与几何结构的实验验证了上述结论及导致不精确比值仍有效的机制。更广泛地,该分解为将基于似然的大语言模型方法迁移至流匹配提供了理论基础,同时区分了精确替换与受控近似。
原文摘要 · Abstract (English)
Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and their old/new differences as log-likelihood ratios. We characterize when these substitutions are valid. For linear Gaussian paths, we exactly decompose endpoint NLL into entropy, a weighted CFM objective, an interior velocity--score residual, and a boundary residual. Thus CFM-only estimates and differences are exact only when the corresponding residuals cancel. At the off-policy population optimum, ordinary CFM is not generally a pointwise NLL estimator, whereas \(w_{\mathrm{sc}}(t)=(1-t)/t\) removes the interior residual; this positive result does not extend generally to training or on-policy alignment. On-policy log-ratios can remain biased even for identical endpoint laws or after surrogate optimization. Experiments across dimensions, distributions, and geometries support these conclusions and the mechanisms that make inexact ratios useful. **More broadly, the decomposition provides a theoretical basis for adapting likelihood-based LLM methods to flow matching, while distinguishing exact substitutions from controlled surrogates.**
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。