arXiv:2506.09988cs.CVcs.AI2025-06ACL被引 3

构建文本编辑评估基准,揭示现有模型在判断图像修改时的缺陷。

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

  • 基于人工标注设计多维度验证模板,构建新评估基准
  • 发现当前模型在识别伪影和描述变化上准确率不足
  • 提出新方法,提升伪影检测与修改描述生成能力

文本引导的图像编辑因生成式AI的发展而日益普及,亟需一套全面的评估框架来验证编辑质量。为此,我们提出EditInspector,一个基于大规模人工标注的文本引导图像编辑评估基准,采用详尽的编辑验证模板收集数据。利用该基准,我们评估了当前最先进(SoTA)视觉语言模型在准确性、伪影检测、视觉质量、与场景融合度、常识符合性以及修改描述生成等方面的性能。结果表明,现有模型在综合评估中表现不佳,且常在描述变化时出现幻觉。针对此问题,我们提出两种新方法,在伪影检测和差异描述生成任务上均优于现有模型。

原文摘要 · Abstract (English)

Text-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their quality. To address this need, we introduce EditInspector, a novel benchmark for evaluation of text-guided image edits, based on human annotations collected using an extensive template for edit verification. We leverage EditInspector to evaluate the performance of state-of-the-art (SoTA) vision and language models in assessing edits across various dimensions, including accuracy, artifact detection, visual quality, seamless integration with the image scene, adherence to common sense, and the ability to describe edit-induced changes. Our findings indicate that current models struggle to evaluate edits comprehensively and frequently hallucinate when describing the changes. To address these challenges, we propose two novel methods that outperform SoTA models in both artifact detection and difference caption generation.

图像编辑评估基准生成模型文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。