用改进的强化学习方法,让AI生成图像视频更符合人类偏好。
DanceGRPO: Unleashing GRPO on Visual Generation
- 将GRPO算法改进后用于视觉生成,提升训练稳定性。
- 在多个模型和任务上表现优异,最高性能提升181%。
- 适合需要高质量图像视频生成的研究者与开发者。
生成式AI已极大推动视觉内容创作,但如何使模型输出符合人类偏好仍是关键挑战。尽管强化学习(RL)被视为优化生成模型的有效途径,但现有方法如DDPO和DPOK在扩展至大规模、多样化提示集时难以维持稳定优化,严重限制其实用性。本文提出DanceGRPO,通过创新性地将组相对策略优化(GRPO)适配于视觉生成任务,利用其内在稳定性机制克服此前基于RL方法在视觉生成中的优化难题。DanceGRPO实现多项突破:其一,在扩散模型与修正流等多种生成范式中均实现稳定策略优化;其二,可在包含三大任务与四类基础模型的复杂真实场景中保持稳健性能;其三,能有效优化五种不同奖励模型所捕获的人类偏好,涵盖图像/视频美学、图文对齐、视频运动质量及二元反馈。全面实验表明,DanceGRPO在多个基准测试(包括HPS-v2.1、CLIP Score、VideoAlign、GenEval)中相较基线方法最高提升达181%。结果证明DanceGRPO是视觉生成中可扩展的强化学习从人类反馈(RLHF)任务的可靠且通用解决方案,为强化学习与视觉合成的融合提供新洞见。
原文摘要 · Abstract (English)
Recent advances in generative AI have revolutionized visual content creation, yet aligning model outputs with human preferences remains a critical challenge. While Reinforcement Learning (RL) has emerged as a promising approach for fine-tuning generative models, existing methods like DDPO and DPOK face fundamental limitations - particularly their inability to maintain stable optimization when scaling to large and diverse prompt sets, severely restricting their practical utility. This paper presents DanceGRPO, a framework that addresses these limitations through an innovative adaptation of Group Relative Policy Optimization (GRPO) for visual generation tasks. Our key insight is that GRPO's inherent stability mechanisms uniquely position it to overcome the optimization challenges that plague prior RL-based approaches on visual generation. DanceGRPO establishes several significant advances: First, it demonstrates consistent and stable policy optimization across multiple modern generative paradigms, including both diffusion models and rectified flows. Second, it maintains robust performance when scaling to complex, real-world scenarios encompassing three key tasks and four foundation models. Third, it shows remarkable versatility in optimizing for diverse human preferences as captured by five distinct reward models assessing image/video aesthetics, text-image alignment, video motion quality, and binary feedback. Our comprehensive experiments reveal that DanceGRPO outperforms baseline methods by up to 181\% across multiple established benchmarks, including HPS-v2.1, CLIP Score, VideoAlign, and GenEval. Our results establish DanceGRPO as a robust and versatile solution for scaling Reinforcement Learning from Human Feedback (RLHF) tasks in visual generation, offering new insights into harmonizing reinforcement learning and visual synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。