打造高精度奖励模型,让图像编辑强化学习真正跑起来。
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
- 构建专用奖励模型EditScore,匹配大模型性能。
- 在基准测试中,最大版本超越GPT-5表现。
- 适合希望用强化学习提升图像编辑能力的研究者。
指令引导的图像编辑已取得显著进展,但面对复杂指令仍需多次尝试才能达成目标。强化学习(RL)有望解决此问题,但其应用受限于缺乏高质量、高效的奖励信号。本文提出系统性方法,核心是开发前沿的专用奖励模型。首先构建EditReward-Bench,用于系统评估编辑质量的奖励模型。基于此,我们推出一系列从7B到72B参数的EditScore模型。通过精细数据筛选,EditScore性能可媲美训练有素的私有视觉语言模型。结合针对生成任务优化的自集成策略,最大版本在基准上甚至超越GPT-5。实验表明,高保真奖励模型是实现在线强化学习的关键。尽管现有最大开源视觉语言模型无法提供有效学习信号,EditScore却能实现高效稳健的策略优化。将该框架应用于强基座模型OmniGen2,最终模型获得显著且一致的性能提升。本工作首次系统性地打通从基准评测、奖励建模到强化学习训练的全流程,证明领域专用的高保真奖励模型是释放强化学习在图像编辑中潜力的核心。
原文摘要 · Abstract (English)
Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning (RL) offers a promising solution, but its adoption in image editing has been severely hindered by the lack of a high-fidelity, efficient reward signal. In this work, we present a comprehensive methodology to overcome this barrier, centered on the development of a state-of-the-art, specialized reward model. We first introduce EditReward-Bench, a comprehensive benchmark to systematically evaluate reward models on editing quality. Building on this benchmark, we develop EditScore, a series of reward models (7B-72B) for evaluating the quality of instruction-guided image editing. Through meticulous data curation and filtering, EditScore effectively matches the performance of learning proprietary VLMs. Furthermore, coupled with an effective self-ensemble strategy tailored for the generative nature of EditScore, our largest variant even surpasses GPT-5 in the benchmark. We then demonstrate that a high-fidelity reward model is the key to unlocking online RL for image editing. Our experiments show that, while even the largest open-source VLMs fail to provide an effective learning signal, EditScore enables efficient and robust policy optimization. Applying our framework to a strong base model, OmniGen2, results in a final model that shows a substantial and consistent performance uplift. Overall, this work provides the first systematic path from benchmarking to reward modeling to RL training in image editing, showing that a high-fidelity, domain-specialized reward model is the key to unlocking the full potential of RL in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。