用强化学习提升图像生成与编辑质量,显著改善对提示词的理解和视觉效果。
Qwen-Image-2.0-RL Technical Report

- 通过人类反馈强化学习与策略蒸馏,优化图像生成与编辑能力。
- 在多个评测中得分提升,文本生成与编辑的Elo评分分别增加78和93。
- 适合关注高质量图像生成、精准指令跟随的开发者与研究者。
我们提出Qwen-Image-2.0-RL,一个后训练流程,通过人类反馈强化学习(RLHF)与在线策略蒸馏(OPD)提升Qwen-Image-2.0扩散模型的视觉质量与指令遵循能力。为提供可靠奖励信号,我们基于点对点评分范式与思维链推理,微调视觉语言模型构建任务专用复合奖励模型:文本到图像生成涵盖对齐性、美学与人像保真度;图像编辑任务则聚焦指令准确性和人脸身份保留。在此基础上,我们设计了基于GRPO的可扩展强化学习框架,包含混合无分类器指引(CFG)策略以保持预训练知识,通过组内奖励范围过滤进行提示筛选,以及按类别校准奖励权重。为融合文本生成与编辑的专用策略,我们提出在线策略蒸馏,通过轨迹级速度匹配将多教师模型合并为单一学生模型。大量评估显示,该模型在Qwen-Image-Bench上取得57.84分(较基线提升2.61),文本生成与图像编辑的Elo评分分别为1193(+78)和1349(+93),在美学质量、提示遵循与编辑准确性方面均实现持续提升。
原文摘要 · Abstract (English)
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instruction-following capability of the Qwen-Image-2.0 diffusion model. To provide reliable reward signals, we construct task-specific composite reward models by fine-tuning vision-language models with a pointwise scoring paradigm and chain-of-thought reasoning. For text-to-image generation, the reward models cover alignment, aesthetics, and portrait fidelity dimensions. For image editing tasks, the reward system addresses instruction-following accuracy and face identity preservation. Building on this reward system, we develop a scalable GRPO-based RL training framework, incorporating a hybrid classifier-free guidance (CFG) strategy to preserve pre-trained knowledge, prompt curation via intra-group reward range filtering, and per-category reward weight calibration. To merge the task-specialized RL policies for T2I and editing, we propose on-policy distillation as the final training stage, which consolidates multiple teachers into a single student model through trajectory-level velocity matching. Extensive evaluation shows that Qwen-Image-2.0-RL achieves 57.84 overall score on Qwen-Image-Bench (+2.61 over the base model), Elo ratings of 1193 in text-to-image arena (+78) and 1349 in image edit arena (+93), demonstrating consistent gains in aesthetic quality, prompt adherence, and editing accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。