arXiv:2507.16193cs.CVcs.MM2025-07被引 23

首个18K规模图像编辑评估基准,用大模型提升编辑质量评测准确性。

LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

  • 构建18K图像编辑数据集,含细粒度人类偏好标注。
  • 提出LMM4Edit统一评估编辑质量、对齐性与属性保留,媲美人评。
  • 支持零样本泛化,适用于多种编辑任务评测。

文本引导图像编辑(TIE)技术快速发展,但现有模型在图像质量、编辑对齐性和原始图像一致性之间难以平衡,限制了实际应用。现有评估基准和指标在规模或与人类感知对齐方面存在局限。为此,我们提出EBench-18K,首个大规模图像编辑基准,包含18,000+张编辑图像,1,080个源图像,21类任务,17种先进TIE模型生成的18,000+张编辑图,55,000+个均值意见评分(MOS),以及18,000+个问答对。基于该数据集,我们利用先进大模型(LMM)评估编辑结果,并分析其理解能力与人类偏好的一致性。进一步提出LMM4Edit,一种基于大模型的统一评估指标,可同时衡量感知质量、编辑对齐性、属性保持及任务特定问答准确率。大量实验表明,LMM4Edit表现优异且与人类偏好高度一致。在其他数据集上的零样本验证也展示了其良好的泛化能力。数据集与代码已开源:https://github.com/IntMeGroup/LMM4Edit。

原文摘要 · Abstract (English)

The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.

图像编辑大模型评估多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。