梳理随机最优控制损失函数分类,揭示其梯度本质相同但方差不同。
A Taxonomy of Loss Functions for Stochastic Optimal Control
- 按期望梯度相同性将损失函数分组,统一优化景观
- 实验表明不同损失函数在梯度方差上差异显著影响性能
- 适合研究扩散模型、流匹配等生成模型的训练机制
随机最优控制(SOC)旨在调控噪声系统,广泛应用于科学、工程和人工智能领域。扩散模型与流匹配模型的奖励微调及无归一化采样方法均可转化为SOC问题。近期工作引入伴随匹配(Adjoint Matching,Domingo-Enrich et al., 2024),在奖励微调设置中显著优于现有损失函数。本文旨在厘清所有现有(及部分新)SOC损失函数之间的联系。我们证明,这些损失函数可按其期望梯度一致性划分为若干类别,这意味着它们具有相同的优化景观,仅在梯度方差上存在差异。通过简单的SOC实验,我们分析了不同损失函数的优势与局限性。
原文摘要 · Abstract (English)
Stochastic optimal control (SOC) aims to direct the behavior of noisy systems and has widespread applications in science, engineering, and artificial intelligence. In particular, reward fine-tuning of diffusion and flow matching models and sampling from unnormalized methods can be recast as SOC problems. A recent work has introduced Adjoint Matching (Domingo-Enrich et al., 2024), a loss function for SOC problems that vastly outperforms existing loss functions in the reward fine-tuning setup. The goal of this work is to clarify the connections between all the existing (and some new) SOC loss functions. Namely, we show that SOC loss functions can be grouped into classes that share the same gradient in expectation, which means that their optimization landscape is the same; they only differ in their gradient variance. We perform simple SOC experiments to understand the strengths and weaknesses of different loss functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。