arXiv:2509.26346cs.CVcs.AI2025-09中稿 · ICLR被引 53

用专家标注数据训练的奖励模型,提升图像编辑指令对齐效果。

EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing

  • 基于20万+人类偏好对构建专家级标注数据集,训练对齐人类偏好的奖励模型。
  • 在GenAI-Bench等4个基准上超越现有VLM类裁判模型,人机相关性达最优。
  • 可筛选高质量数据,助力开源图像编辑模型训练,适合研究者与开发者使用。

近期自然语言指令驱动的图像编辑取得显著进展,但开源模型仍落后于闭源模型。主要瓶颈在于缺乏可靠的奖励模型以生成高质量合成训练数据。为此,我们构建了EditReward,基于超过20万条由专业人员严格标注的人类偏好对数据集进行训练。实验表明,EditReward在指令引导图像编辑任务中展现出更优的人类偏好对齐能力,在GenAI-Bench、AURORA-Bench、ImagenHub及新提出的EditReward-Bench等多个基准上均达到当前最佳的人机相关性表现,优于多种VLM-as-judge模型。此外,我们利用EditReward从噪声数据集ShareGPT-4o-Image中筛选高质量子集,并在此基础上训练Step1X-Edit模型,其性能显著优于使用完整数据集训练的结果,验证了EditReward在提升训练数据质量方面的有效性。其强对齐特性也预示其在强化学习后训练与测试时扩展等高级应用中的潜力。EditReward及其训练数据将公开发布,助力社区构建更高质量的图像编辑训练数据集。

原文摘要 · Abstract (English)

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. However, the open-source models are still lagging. The main bottleneck is the lack of a reliable reward model to scale up high-quality synthetic training data. To address this critical bottleneck, we built EditReward, trained with our new large-scale human preference dataset, meticulously annotated by trained experts following a rigorous protocol containing over 200K preference pairs. EditReward demonstrates superior alignment with human preferences in instruction-guided image editing tasks. Experiments show that EditReward achieves state-of-the-art human correlation on established benchmarks such as GenAI-Bench, AURORA-Bench, ImagenHub, and our new EditReward-Bench, outperforming a wide range of VLM-as-judge models. Furthermore, we use EditReward to select a high-quality subset from the existing noisy ShareGPT-4o-Image dataset. We train Step1X-Edit on the selected subset, which shows significant improvement over training on the full set. This demonstrates EditReward's ability to serve as a reward model to scale up high-quality training data for image editing. Furthermore, its strong alignment suggests potential for advanced applications like reinforcement learning-based post-training and test-time scaling of image editing models. EditReward with its training dataset will be released to help the community build more high-quality image editing training datasets.

图像编辑奖励模型人类对齐数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。