arXiv:2606.15570cs.CV2026-06IJCV

首个兼顾单轮与多轮指令图像编辑的全面评估基准。

An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing

论文配图:An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing
图 1 · 摘自论文原文
  • 构建涵盖16维单轮和7维多轮的评估体系,覆盖高阶与低阶指标。
  • 通过用户研究确保评估标准与人类判断高度一致。
  • 为当前主流编辑模型提供定位短板的实证分析,指导未来研究。

近年来,基于指令的图像编辑(IIE)取得了显著进展,但其评估仍面临挑战,主要源于指令的复杂性和编辑类型的多样性。为解决这一问题,本文提出一个全面的评估基准I2EBench2.0,支持单轮与多轮IIE模型的评估。该基准具备四大特点:1)同时评估单轮与多轮编辑任务,衡量编辑的精确性与一致性;2)包含16个单轮维度和7个多轮维度的丰富评估标准,覆盖高低层特征;3)通过大规模用户研究确保评估标准与人类判断对齐;4)通过对8个前沿IIE模型的系统测试,揭示其在各维度上的优劣表现,为未来研究提供关键洞见。相关代码、数据集及生成图像均已开源于GitHub。

原文摘要 · Abstract (English)

In recent years, there have been notable advancements in the area of instruction-based image editing (IIE), which focuses on the automatic alteration of input images using a model. Nevertheless, assessing the effectiveness of these editing models poses a considerable challenge due to the intricate nature of instructions and the wide variety of edits. To tackle this problem, one urgent task in this domain is the development of a robust evaluation framework that can precisely gauge the quality of editing outcomes and offer valuable benchmarks to guide future improvements. To address this challenge, we present a comprehensive evaluation benchmark named I2EBench2.0, designed for single-round and multi-round assessment of IIE models. I2EBench2.0 has four key features: 1) Evaluation Across Single and Multi-rounds: I2EBench2.0 simultaneously evaluates both single-round and multi-round instruction-based edits, assessing the precision and consistency of the edits. 2) Extensive Evaluation Criteria: I2EBench2.0 encompasses a broad range of criteria, evaluating both high-level and low-level aspects of each IIE model. Specifically, it incorporates 16 dimensions for single-round evaluations and 7 for multi-round evaluations. 3) Alignment with Human Judgment: To ensure our benchmark aligns with human evaluation, we conducted a comprehensive user study for each criterion. 4) Research-driven Insights: By analyzing the strengths and weaknesses of current IIE models across all 16 single-round and 7 multi-round dimensions, we provide critical insights aimed at directing future research in this area. We tested eight recently developed IIE models using I2EBench2.0 and derived academic insights through meticulous comparison and analysis. The related code, dataset, and images generated by all IIE models are available on GitHub: https://github.com/cocoshe/I2EBench.

图像编辑评估基准指令理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。