提出新方法提升生成模型对人类偏好的响应能力。
Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
- 将奖励微调建模为无记忆随机最优控制问题。
- 新算法在真实感、一致性与泛化性上超越现有方法。
- 适合需要精准控制生成结果的场景,如艺术创作。
通过迭代过程生成样本的动力学生成模型(如流匹配和去噪扩散模型)应用广泛,但缺乏理论严谨的奖励微调方法。本文将奖励微调建模为随机最优控制(SOC),并证明微调时必须采用特定的无记忆噪声调度,以处理噪声变量与生成样本间的依赖关系。为此提出新算法Adjoint Matching,将SOC问题转化为回归任务,显著优于现有方法,在真实感、一致性和未见人类偏好奖励模型上的泛化能力方面表现更优,同时保持样本多样性。
原文摘要 · Abstract (English)
Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fine-tuning as stochastic optimal control (SOC). Critically, we prove that a very specific memoryless noise schedule must be enforced during fine-tuning, in order to account for the dependency between the noise variable and the generated samples. We also propose a new algorithm named Adjoint Matching which outperforms existing SOC algorithms, by casting SOC problems as a regression problem. We find that our approach significantly improves over existing methods for reward fine-tuning, achieving better consistency, realism, and generalization to unseen human preference reward models, while retaining sample diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。