arXiv:2605.13062cs.CV2026-05被引 2

构建统一评测基准,精准评估图像编辑与奖励模型性能

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

论文配图:Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
图 1 · 摘自论文原文
  • 设计六类渐进难题,覆盖跨图像推理与多图编辑能力
  • 2388个标注样本+细粒度评分体系,更贴近人类判断
  • 配套2251组偏好数据,模拟真实强化学习优化场景

近期图像编辑模型在指令遵循、多模态理解及复杂视觉编辑方面取得显著进展,但现有评测基准常因任务难度不足和粗粒度评价方式,难以真实反映人类判断,尤其对前沿强模型表现不佳。同时,奖励模型在基于强化学习的图像编辑优化中日益关键,但现有奖励模型评测仍依赖脱离实际应用的设定。为解决这些问题,我们提出 Edit-Compass 和 EditReward-Compass,一套统一的图像编辑与奖励建模评测套件。Edit-Compass 包含 2,388 个精心标注的实例,涵盖六类逐步增加难度的任务类别,涵盖世界知识推理、视觉推理及多图像编辑等能力。除广泛任务覆盖外,该基准采用基于结构化推理的细粒度多维度评价框架与精心设计的评分标准。同时,EditReward-Compass 提供 2,251 组偏好对,模拟强化学习优化过程中的真实奖励建模场景。

原文摘要 · Abstract (English)

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.

图像编辑评测基准强化学习奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。