解决均值流训练中梯度方差过大的问题,给出理论解释和优化方案。
On Variance Reduction in Learning Mean Flows

- 发现条件速度场在损失函数中承担双重统计角色,原方法系数设定错误。
- 推导出最优系数闭式解,在二维与潜空间扩散变换器上验证其效果。
- 证明方差最优系数与生成质量最优系数不一致,解释了现有方法的差异。
单步生成建模已成为降低扩散模型与流匹配模型推理成本的主流方法。在无蒸馏方法中,均值流(MeanFlow)训练长期存在不稳定问题,表现为损失不下降、梯度方差无界。本文建立理论,指出该病态源于对条件速度场的误用。我们揭示条件速度在损失中同时扮演无偏回归目标与雅可比-向量积中的蒙特卡洛控制变量双重角色,而原均值流损失为后者分配了错误系数。我们推导出最优系数的闭式表达,并表明近期多种修正方法实为同一最优解的不同实现。在二维基准与潜空间扩散变压器(DiT)上的可控参数扫描验证了预测的偏差-方差权衡。其中DiT实验揭示了量化FID与均方误差(MSE)之间的景观错配:尽管梯度MSE在β≈0.94的内部系数处最小化,但最低FID却偏好直接使用条件速度的无偏端点。本分析解释了均值流的不稳定性,统一了现有修正方案,并表明方差最优系数未必对应质量最优。
原文摘要 · Abstract (English)
One-step generative modeling has emerged as a leading approach for amortizing the inference cost of diffusion and flow-matching models. Among distillation-free methods, MeanFlow training is notoriously unstable, with non-decreasing loss and unbounded gradient variance. In this work, we establish a theory that attributes this pathology to a misuse of the conditional velocity field. We show that the conditional velocity plays two distinct statistical roles in the loss: both as an unbiased regression target and as a Monte Carlo control variate in a Jacobi-vector product, with the original MeanFlow loss assigning the wrong coefficient to the latter. We derive the optimal coefficient in closed form and show that a family of fixes in concurrent works corresponds to different practical realizations of the same optimum. A controlled sweep of this coefficient on two-dimensional benchmarks and on a latent Diffusion Transformer recovers the predicted bias-variance ordering. Our DiT experiment also reveals a quantitative FID-MSE landscape mismatch. Specifically, although the gradient-MSE is minimized at an interior coefficient value near $β\!=\!0.94$, the coefficient that minimizes FID prefers to use conditional velocity directly at the unbiased corner. Our analysis therefore explains why MeanFlow is unstable and unifies its concurrent remedies, and shows that the variance-optimal coefficient need not coincide with the quality-optimal one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。