构建首个大规模图文编辑评估基准,用大模型探针提升评测准确性。
Evaluating Image Editing with LLMs: A Comprehensive Benchmark and Intermediate-Layer Probing Approach
- 设计TIEdit基准,覆盖8类任务、5120张编辑图像,由专家打分生成1.5万条主观评分。
- 提出EditProbe方法,通过多模态大模型中间层特征捕捉语义与感知关系。
- 实验证明现有自动指标相关性弱,而新方法更贴近人类判断,适合模型开发者使用。
文本引导的图像编辑(TIE)评估仍具挑战性,需兼顾视觉质量、指令对齐和内容保留。尽管TIE模型进展迅速,现有评估基准规模有限且与人类感知相关性弱。本文提出TIEdit,一个系统性评估框架:包含512张源图、8类典型编辑任务、10个先进TIE模型生成的5120张编辑图像。招募20名专家进行307,200次主观评分,汇总得15,360个均值意见分数(MOS),覆盖视觉质量、编辑对齐、内容保留三个维度。进一步提出EditProbe,基于大语言模型中间层探针的评估方法,通过提取多模态大模型中间表示,捕捉源图、指令与结果间的语义与感知关系。实验表明,现有自动指标与人类判断相关性低,而EditProbe显著提升一致性。TIEdit与EditProbe共同为更可靠、感知一致的TIE评估提供基础。
原文摘要 · Abstract (English)
Evaluating text-guided image editing (TIE) methods remains a challenging problem, as reliable assessment should simultaneously consider perceptual quality, alignment with textual instructions, and preservation of original image content. Despite rapid progress in TIE models, existing evaluation benchmarks remain limited in scale and often show weak correlation with human perceptual judgments. In this work, we introduce TIEdit, a benchmark for systematic evaluation of text-guided image editing methods. TIEdit consists of 512 source images paired with editing prompts across eight representative editing tasks, producing 5,120 edited images generated by ten state-of-the-art TIE models. To obtain reliable subjective ratings, 20 experts are recruited to produce 307,200 raw subjective ratings, which accumulates into 15,360 mean opinion scores (MOSs) across three evaluation dimensions: perceptual quality, editing alignment, and content preservation. Beyond the benchmark itself, we further propose EditProbe, an LLM-based evaluator that estimates editing quality via intermediate-layer probing of hidden representations. Instead of relying solely on final model outputs, EditProbe extracts informative representations from intermediate layers of multimodal large language models to better capture semantic and perceptual relationships between source images, editing instructions, and edited results. Experimental results demonstrate that widely used automatic evaluation metrics show limited correlation with human judgments on editing tasks, while EditProbe achieves substantially stronger alignment with human perception. Together, TIEdit and EditProbe provide a foundation for more reliable and perceptually aligned evaluation of text-guided image editing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。