arXiv:2409.08400cs.LGcs.AI2024-09被引 20

用连续时间强化学习优化扩散模型生成质量,让评分变成可调控的行动。

Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning

  • 将得分函数视为控制动作,构建连续时间强化学习框架
  • 在文本到图像生成任务中显著提升生成质量与对齐度
  • 适合研究生成模型对齐与强化学习融合的学者

从人类反馈中学习奖励函数来微调扩散生成模型的任务,被严格地形式化为一个探索性的连续时间随机控制问题。本文的核心思想是将得分匹配函数视为控制动作,并在此基础上,从连续时间视角发展出一个统一的框架,利用强化学习算法提升扩散模型的生成质量。同时,在随机微分方程驱动环境的假设下,建立了相应的连续时间强化学习理论,用于策略优化与正则化。实验将在附带论文中报告,聚焦于文本到图像(T2I)生成任务。

原文摘要 · Abstract (English)

Reinforcement Learning from human feedback (RLHF) has been shown a promising direction for aligning generative models with human intent and has also been explored in recent works for alignment of diffusion generative models. In this work, we provide a rigorous treatment by formulating the task of fine-tuning diffusion models, with reward functions learned from human feedback, as an exploratory continuous-time stochastic control problem. Our key idea lies in treating the score-matching functions as controls/actions, and upon this, we develop a unified framework from a continuous-time perspective, to employ reinforcement learning (RL) algorithms in terms of improving the generation quality of diffusion models. We also develop the corresponding continuous-time RL theory for policy optimization and regularization under assumptions of stochastic different equations driven environment. Experiments on the text-to-image (T2I) generation will be reported in the accompanied paper.

扩散模型强化学习生成对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。