针对流模型强化学习中的能力错配问题,提出自适应增强方法提升对齐效果。
AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

- 根据模型当前能力动态筛选提示词,实现分层学习
- 融合组内与全局优势评估,更准确衡量策略改进
- 可无缝接入现有框架,稳定训练并持续提效
基于流的组相对策略优化(GRPO)在对齐文本到图像生成模型与人类偏好方面表现卓越。然而,我们发现现有方法在学习循环中与学习者当前能力脱节,存在提示词选择和优势估计的双重盲区:(i) 当前方法随机采样提示词,忽视数据选择对强化学习效率的关键影响——该因素在大语言模型的GRPO中已被证明至关重要;(ii) 仅依赖组内统计评估样本质量,缺乏全局视角,难以准确衡量真实策略进步。为此,我们提出能力感知的自适应强化学习算法AdaGRPO,包含两个核心组件:(i) 在线课程过滤策略:动态追踪模型能力,自适应选择最匹配其当前学习边界的提示词;(ii) 跨层级优势融合:协同整合细粒度组内优势与宏观级全局优势,实现全面且无偏的策略评估。作为轻量级、即插即用模块,AdaGRPO可无缝集成至Flow-GRPO、DanceGRPO和Flow-CPS等现有框架。大量实验表明,AdaGRPO能持续提升性能,并显著稳定流模型的GRPO训练。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that the learning loop of current flow-based GRPO is fundamentally decoupled from the learner's current capability, suffering from critical blind spots at both prompt selection and advantage estimation: (i) Existing methods sample prompts randomly, overlooking the substantial impact of data selection on reinforcement learning (RL) efficacy--a factor proven crucial in GRPO for large language models; (ii) They evaluate sample quality solely relying on intra-group statistics, lacking a global perspective to accurately measure true policy improvement. To address these issues, we propose Adaptive GRPO (AdaGRPO), a novel capability-aware RL algorithm tailored for flow models. Specifically, AdaGRPO consists of two principal components: (i) Online Curriculum Filtering Strategy: Dynamically tracks the model's proficiency and adaptively selects prompts that best match its current learning boundary; (ii) Cross-Level Advantage Fusion: Synergistically integrates fine-grained intra-group advantages with macro-level global advantages, providing a comprehensive and unbiased policy evaluation. As a lightweight, plug-and-play module, AdaGRPO can be seamlessly integrated with existing frameworks such as Flow-GRPO, DanceGRPO, and Flow-CPS. Extensive experiments demonstrate that AdaGRPO consistently drives performance gains while significantly stabilizes GRPO training for flow models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。