用轨迹流优化提升文生图模型的生成质量与提示对齐
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
- 将采样过程视为单一动作,减少更新方差
- 收敛更快,生成质量与提示匹配度更高
- 适合追求高质量图像生成的研究者与开发者
强化学习已成为扩散模型后训练的标准方法,可利用奖励信号显式提升图像质量与提示对齐。本文提出一种在线强化学习变体,通过采样成对轨迹并沿更优图像方向调整采样流速,降低模型更新方差。不同于将每步采样视为独立策略动作的方法,本文将整个采样过程视为一个动作。实验使用高质量视觉语言模型和现成质量指标作为奖励,并在多种评估指标下验证效果。结果表明,该方法收敛更快,生成图像质量更高,提示对齐性更强,优于以往方法。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment. In this paper, we propose an online RL variant that reduces the variance in the model updates by sampling paired trajectories and pulling the flow velocity in the direction of the more favorable image. Unlike existing methods that treat each sampling step as a separate policy action, we consider the entire sampling process as a single action. We experiment with both high-quality vision language models and off-the-shelf quality metrics for rewards, and evaluate the outputs using a broad set of metrics. Our method converges faster and yields higher output quality and prompt alignment than previous approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。