新基准GEditBench v2提升图像编辑评估真实感,更贴近人类判断。
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
- 构建1200个真实用户指令的多任务评测集,含开放域编辑任务
- 提出PVC-Judge模型,对图像编辑一致性评估达领先水平
- 适合研究图像编辑、评估方法及人机对齐的学者使用
近期图像编辑技术已能实现复杂指令下的高保真生成,但现有评估框架存在任务覆盖窄、视觉一致性度量不足等问题。为此,我们推出GEditBench v2,包含1200个真实用户查询,涵盖23类任务,新增开放集以支持未预定义的编辑指令。我们提出PVC-Judge,一个开源的成对评估模型,通过两种新型区域解耦偏好数据合成管道训练。同时构建VCReward-Bench,基于专家标注的偏好对,评估PVC-Judge与人类判断的一致性。实验表明,PVC-Judge在开源模型中表现最优,甚至超越GPT-5.1平均表现。对16个前沿编辑模型的测评显示,GEditBench v2可更准确揭示模型缺陷,为精准图像编辑提供可靠评估基础。
原文摘要 · Abstract (English)
Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks suffer from narrow task coverage, while standard metrics fail to adequately capture visual consistency, i.e., the preservation of identity, structure and semantic coherence between edited and original images. To address these limitations, we introduce GEditBench v2, a comprehensive benchmark with 1,200 real-world user queries spanning 23 tasks, including a dedicated open-set category for unconstrained, out-of-distribution editing instructions beyond predefined tasks. Furthermore, we propose PVC-Judge, an open-source pairwise assessment model for visual consistency, trained via two novel region-decoupled preference data synthesis pipelines. Besides, we construct VCReward-Bench using expert-annotated preference pairs to assess the alignment of PVC-Judge with human judgments on visual consistency evaluation. Experiments show that our PVC-Judge achieves state-of-the-art evaluation performance among open-source models and even surpasses GPT-5.1 on average. Finally, by benchmarking 16 frontier editing models, we show that GEditBench v2 enables more human-aligned evaluation, revealing critical limitations of current models, and providing a reliable foundation for advancing precise image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。