揭示流模型在确定性训练下的记忆失效机制,发现路径直线化会导致错误配对。
Gradient Variance Reveals Failure Modes in Flow-Based Generative Models
- 通过梯度方差分析,发现直线路径目标会诱导模型记忆训练配对。
- 在所有插值线相交的场景下,推理时仍复现训练配对,导致泛化失败。
- 加入微小噪声可恢复泛化能力,适用于生成模型可靠性研究。
修正流模型学习常微分方程向量场,使源分布与目标分布间的轨迹为直线,实现近似一步推理。我们发现,这种直线路径目标掩盖了根本性失效模式:在确定性训练下,低梯度方差驱动模型记忆任意训练配对,即使配对之间的插值线相交。通过研究高斯到高斯的传输问题,利用随机与确定性设置下的损失梯度方差,刻画优化在不同情形下偏好何种向量场。我们进一步证明,在所有插值线相交的情况下,应用修正流模型将产生与训练时相同的特定配对;更一般地,即使训练插值线相交,也存在一个记忆型向量场,且优化直线路径目标会收敛到该病态场。推理阶段采用确定性积分,会精确重现训练配对。我们在CelebA数据集上验证了这一现象,确认确定性插值引发记忆,而引入微小噪声可恢复泛化能力。
原文摘要 · Abstract (English)
Rectified Flows learn ODE vector fields whose trajectories are straight between source and target distributions, enabling near one-step inference. We show that this straight-path objective conceals fundamental failure modes: under deterministic training, low gradient variance drives memorization of arbitrary training pairings, even when interpolant lines between pairs intersect. To analyze this mechanism, we study Gaussian-to-Gaussian transport and use the loss gradient variance across stochastic and deterministic regimes to characterize which vector fields optimization favors in each setting. We then show that, in a setting where all interpolating lines intersect, applying Rectified Flow yields the same specific pairings at inference as during training. More generally, we prove that a memorizing vector field exists even when training interpolants intersect, and that optimizing the straight-path objective converges to this ill-defined field. At inference, deterministic integration reproduces the exact training pairings. We validate our findings empirically on the CelebA dataset, confirming that deterministic interpolants induce memorization, while the injection of small noise restores generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。