提出面向多轮图像编辑的细粒度自动化评估框架,解决现有方法依赖参考图或提示不准确的问题。
EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- 基于对象中心视角,分解图像并动态更新对象池以支持多轮评估
- 设计三项新指标:指令遵循度、内容一致性与视觉质量,均在真实编辑流程中验证
- 构建涵盖16个模型的多轮编辑基准,助力发现模型缺陷并指导改进
基于指令的图像编辑发展迅速,但可靠且可解释的评估仍是瓶颈。现有方法要么依赖成对参考图,覆盖有限且继承生成模型偏差;要么仅依赖零样本视觉语言模型(VLMs),其通过提示评估指令遵循性、内容一致性和视觉质量往往不够精确。为此,我们提出EdiVal,一个基于对象中心视角的自动化、细粒度评估框架,可精准评估单轮及多轮指令式编辑。给定输入图像,EdiVal首先将其分解为语义有意义的对象,再合成多样化上下文感知的编辑指令,并在多轮中动态更新对象池。该流程催生三项新指标:1)EdiVal-IF,结合开放词汇检测器进行符号检查与VLM在检测引导裁剪上的语义验证,衡量指令遵循度;2)EdiVal-CC,利用演化对象池计算未更改对象与背景的语义相似度,评估内容一致性;3)EdiVal-VQ,借助人类偏好模型量化整体视觉质量变化。我们构建了EdiVal Bench,一个覆盖9类指令和16个先进编辑模型的多轮编辑基准,涵盖上下文、流匹配与扩散范式。实验证明,EdiVal可识别现有模型的失效模式,为下一代编辑模型开发提供依据。
原文摘要 · Abstract (English)
Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images, resulting in limited coverage and inheriting biases from prior generative models or (ii) rely solely on zero-shot vision language models (VLMs), whose prompt-based assessments of instruction following, content consistency, and visual quality are often imprecise. To address this, we introduce EdiVal, an automated and fine-grained evaluation framework grounded in an object-centric perspective, designed to assess not only standard single-turn but also multi-turn instruction-based editing with precision. Given an input image, EdiVal first decomposes it into semantically meaningful objects, then synthesizes diverse, context-aware editing instructions while dynamically updating object pools across turns. These two stages enable two novel object centric metrics tailored for multi turn evaluation and one global metric of visual quality: 1) EdiVal-IF, which measures instruction following by combining open vocabulary object detectors for symbolic checks with VLMs for semantic verification on detector guided crops; 2) EdiVal-CC, which evaluates content consistency by calculating semantic similarity of unchanged objects and background using the evolving object pools; and 3) EdiVal-VQ, which quantifies changes in overall visual quality with human preference models. Instantiating this pipeline, we build EdiVal Bench, a multi-turn editing benchmark covering 9 instruction types and 16 state-of-the-art editing models, spanning in-context, flow-matching, and diffusion paradigms. We demonstrate that EdiVal can be used to identify existing failure modes, thereby informing the development of the next generation of editing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。