arXiv:2505.11493cs.CV2025-05被引 15

新基准GIE-Bench让图文编辑模型评估更精准可靠。

GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing

  • 用多选题自动检验编辑是否准确执行
  • 通过对象感知掩码确保非目标区域不变
  • 适合研究图文编辑的评估与改进

使用自然语言指令编辑图像已成为一种直观且富有表现力的视觉内容修改方式,但评估此类模型性能仍具挑战。现有方法多依赖CLIP等图像-文本相似度指标,精度不足。本文提出GIE-Bench基准,从两个关键维度实现更落地的评估:(i) 功能正确性,通过自动生成的多选题验证目标修改是否成功;(ii) 图像内容保真度,采用对象感知掩码与保真度评分,确保非目标区域保持视觉一致。该基准包含20个内容类别、超1000个高质量编辑样本,每例均配有详细指令、评估问题及空间对象掩码。我们对GPT-Image-1等主流模型进行了大规模对比测试,并将自动指标与人工评分对照验证。结果表明,GPT-Image-1在指令遵循准确性上领先,但常过度修改无关区域,揭示当前模型的核心权衡。GIE-Bench为图文编辑评估提供了可扩展、可复现的框架。

原文摘要 · Abstract (English)

Editing images using natural language instructions has become a natural and expressive way to modify visual content; yet, evaluating the performance of such models remains challenging. Existing evaluation approaches often rely on image-text similarity metrics like CLIP, which lack precision. In this work, we introduce a new benchmark designed to evaluate text-guided image editing models in a more grounded manner, along two critical dimensions: (i) functional correctness, assessed via automatically generated multiple-choice questions that verify whether the intended change was successfully applied; and (ii) image content preservation, which ensures that non-targeted regions of the image remain visually consistent using an object-aware masking technique and preservation scoring. The benchmark includes over 1000 high-quality editing examples across 20 diverse content categories, each annotated with detailed editing instructions, evaluation questions, and spatial object masks. We conduct a large-scale study comparing GPT-Image-1, the latest flagship in the text-guided image editing space, against several state-of-the-art editing models, and validate our automatic metrics against human ratings. Results show that GPT-Image-1 leads in instruction-following accuracy, but often over-modifies irrelevant image regions, highlighting a key trade-off in the current model behavior. GIE-Bench provides a scalable, reproducible framework for advancing more accurate evaluation of text-guided image editing.

图文编辑评估基准图像生成模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。