arXiv:2604.08580math.OCcs.LG2026-04被引 3

用随机最大值原理重新推导了生成模型的对偶匹配方法,让其更严谨可实现。

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control

论文配图:Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control
图 1 · 摘自论文原文
  • 从随机最大值原理出发,统一推导了对偶匹配的目标函数
  • 在扩散项与状态无关时,恢复了简化版损失函数,避免高阶项
  • 揭示了对偶匹配是连续时间逐次逼近法,适合实际训练

扩散模型和流模型的奖励微调,以及从倾斜或玻尔兹曼分布中采样,均可表述为随机最优控制(SOC)问题,其中学习最优生成动力学等价于在随机微分方程约束下优化控制。本文重新审视并推广了近期提出的基于SOC的对偶匹配方法,通过随机最大值原理(SMP)为其提供严格理论基础。我们提出了一个通用的哈密顿对偶匹配目标函数,适用于控制依赖漂移与扩散、且运行成本凸的情况,并证明其期望值与原始SOC目标具有相同的首阶变分。因此,临界点满足哈密顿-雅可比-贝尔曼(HJB)驻定条件。在扩散项与状态和控制均无关的重要情形下,我们恢复了此前提出的简洁对偶匹配损失,该损失规避了二阶项,且其临界点在温和唯一性假设下与最优控制一致。数值实验表明,一旦扩散项依赖状态,这些被忽略的项便变得必要。最后,我们证明对偶匹配可精确解释为由SMP诱导的连续时间逐次逼近法,提供了一种可行的替代方案,克服了传统SMP算法因不可处理鞅项而难以实施的问题。这些结果对随机控制领域亦具独立价值,为实现基于SMP的迭代提供了新路径。

原文摘要 · Abstract (English)

Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints. In this work, we revisit and generalize Adjoint Matching, a recently proposed SOC-based method for learning optimal controls, and place it on a rigorous footing by deriving it from the Stochastic Maximum Principle (SMP). We formulate a general Hamiltonian adjoint matching objective for SOC problems with control-dependent drift and diffusion and convex running costs, and show that its expected value has the same first variation as the original SOC objective. As a consequence, critical points satisfy the Hamilton--Jacobi--Bellman (HJB) stationarity conditions. In the important practical case of state- and control-independent diffusion, we recover the lean adjoint matching loss previously introduced, which avoids second-order terms and whose critical points coincide with the optimal control under mild uniqueness assumptions. Numerical experiments confirm that the extra terms it discards become necessary once the diffusion is state-dependent. Finally, we show that adjoint matching can be precisely interpreted as a continuous-time method of successive approximations induced by the SMP, yielding a practical and implementable alternative to classical SMP-based algorithms, which are obstructed by intractable martingale terms in the stochastic setting. These results are also of independent interest to the stochastic control community, providing new implementable objectives and a viable pathway for SMP-based iterations in stochastic problems.

生成模型最优控制随机最大值原理对偶匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。