首个面向多源图像编辑的人类偏好评估框架,解决现有评测缺失问题。
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

- 基于人类偏好构建多源编辑数据集,覆盖16类任务与36000张生成图。
- 提出MIEScore模型,在视觉质量、指令遵循等维度上逼近人类评分。
- 适合图像编辑研究者、评测工程师及多模态模型开发者使用。
统一多模态模型在文本引导的图像编辑方面取得显著进展,如Nano-Banana-Pro和GPT-Image-2在对象合成、人物背景组合、跨图风格融合等多源图像编辑(MIE)任务中展现出新能力。然而,现有评测基准和图像编辑评估方法仍主要聚焦于单图编辑,严重忽视了更复杂的多源编辑场景。为此,我们提出MIE-Bench,首个大规模多源图像编辑基准,包含3000个编辑实例、16类任务,每项涉及多于两幅源图与编辑提示,并配套12种先进模型生成的36,000张编辑图像,以及超过108,000条平均意见分数(MOS),涵盖视觉质量、指令遵循与属性保持。基于此,我们构建MIEScore——一种基于多模态大语言模型(MLLM)的评估模型,通过技能优化与多维监督微调实现人类对齐反馈。大量实验表明,MIEScore在人类偏好对齐方面达到当前最优性能,且在其他主流评测数据集上具有良好泛化能力。相关数据集与模型已开源:https://github.com/IntMeGroup/MIEScore。
原文摘要 · Abstract (English)
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessment (IEQA) methods remain primarily focused on single-image editing tasks and largely overlook the more challenging setting of MIE. This highlights the urgent need for a comprehensive and human-aligned benchmark for MIE. To this end, we introduce MIE-Bench, the first large-scale multiple image editing benchmark with fine-grained human preference annotations. Specifically, MIE-Bench includes 3,000 editing instances across 16 tasks, each involving more than two source images and an editing prompt, together with 36K edited images produced by 12 state-of-the-art editing models and over 108K mean opinion scores (MOSs) covering visual quality, instruction following, and attribute preservation. Based on MIE-Bench, we propose MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback for MIE. Extensive experiments show that MIEScore achieves state-of-the-art performance in aligning with human preferences and generalizes well across other IEQA datasets. Both the dataset and the model are available at https://github.com/IntMeGroup/MIEScore.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。