arXiv:2509.02528cs.LGmath.OC2025-09被引 2

用偏微分方程方法证明:微调扩散模型可转为更快的回归问题。

Is RL fine-tuning harder than regression? A PDE learning approach for diffusion models

  • 基于哈密顿-雅可比-贝尔曼方程构建变分不等式求解新算法
  • 证明了价值函数与控制策略的精确统计收敛速率
  • 适用于追求高效微调的扩散模型研究者

我们研究利用通用价值函数近似来学习微调给定扩散过程的最优控制策略。通过求解基于哈密顿-雅可比-贝尔曼(HJB)方程的变分不等式问题,开发了一类新算法。我们证明了所学价值函数和控制策略的尖锐统计速率,该速率依赖于函数类的复杂度和近似误差。与通用强化学习问题相比,我们的方法表明微调可通过监督回归实现,且具有更快的统计率保证。

原文摘要 · Abstract (English)

We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the Hamilton-Jacobi-Bellman (HJB) equations. We prove sharp statistical rates for the learned value function and control policy, depending on the complexity and approximation errors of the function class. In contrast to generic reinforcement learning problems, our approach shows that fine-tuning can be achieved via supervised regression, with faster statistical rate guarantees.

扩散模型强化学习微调最优控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。