用大模型当裁判,精细评估图像编辑的12个维度。
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
- 将评价拆解为12个可解释维度,覆盖保真、编辑质量与指令一致性。
- 人评与模型评高度一致,传统指标无法区分过编辑或语义偏差。
- 适合研究者对比和优化图像编辑模型,尤其关注细节控制能力。
图像编辑评估因传统指标粒度粗、可解释性差而困难,常忽略可控性、定位精度和指令忠实度。本文提出基于多模态大模型(MLLM)的细粒度评价框架,将常见评估维度分解为12个细粒度可解释因子,涵盖图像保留、编辑质量与指令一致性。基于此,构建了融合人类判断、MLLM评估、模型输出与传统指标的全新基准,覆盖多种图像编辑任务。通过大规模人类实验验证,所提MLLM裁判在细粒度上与人类评价高度一致,具备可靠性和可扩展性。传统指标常将视觉合理但过编辑或语义模糊的输出误判为优,而本框架在离线与在线场景中均提供更直观、更丰富的评估。本工作提供基准、原则化分解与实证证据,确立细粒度MLLM裁判作为图像编辑研究的基础工具。
原文摘要 · Abstract (English)
Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently reward visually plausible outputs while overlooking controllability, edit localization, and faithfulness to user instructions. In this work, we introduce a fine-grained Multimodal Large Language Model (MLLM)-as-a-Judge framework for image editing that decomposes common evaluation notions into twelve fine-grained interpretable factors spanning image preservation, edit quality, and instruction fidelity. Building on this formulation, we present a new human-validated benchmark that integrates human judgments, MLLM-based evaluations, model outputs, and traditional metrics across diverse image editing tasks. Through extensive human studies, we show that the proposed MLLM judges align closely with human evaluations at a fine granularity, supporting their use as reliable and scalable evaluators. We further demonstrate that traditional image editing metrics are often poor proxies for these factors, failing to distinguish over-edited or semantically imprecise outputs, whereas our judges provide more intuitive and informative assessments in both offline and online settings. Together, this work introduces a benchmark, a principled factorization, and empirical evidence positioning fine-grained MLLM judges as a practical foundation for studying, comparing, and improving image editing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。