arXiv:2508.18159cs.CVcs.LG2025-08被引 2

评测视觉引导图像编辑方法,发现主流模型常误判视觉线索。

SpotEdit: Evaluating Visually-Guided Image Editing Methods

  • 设计SpotEdit基准,覆盖扩散、自回归等多类生成模型
  • 揭示领先模型在视觉线索缺失时仍会错误编辑,存在严重幻觉
  • 适合研究可控生成与评估的AI从业者使用

视觉引导图像编辑通过结合视觉线索和文本提示实现精细可控的内容生成,已成为重要范式。尽管生成模型能力显著提升,现有评估方式仍过于简单,难以反映真实编辑挑战。本文提出SpotEdit,一个全面的基准测试体系,系统评估多种扩散、自回归及混合生成模型在视觉引导编辑任务中的表现,揭示了显著的性能差异。针对一个关键但被忽视的问题,该基准特别包含幻觉检测模块,发现如GPT-4o等领先模型常错误认定视觉线索存在,并据此执行编辑操作。代码与数据集已公开于https://github.com/SaraGhazanfari/SpotEdit。

原文摘要 · Abstract (English)

Visually-guided image editing, where edits are conditioned on both visual cues and textual prompts, has emerged as a powerful paradigm for fine-grained, controllable content generation. Although recent generative models have shown remarkable capabilities, existing evaluations remain simple and insufficiently representative of real-world editing challenges. We present SpotEdit, a comprehensive benchmark designed to systematically assess visually-guided image editing methods across diverse diffusion, autoregressive, and hybrid generative models, uncovering substantial performance disparities. To address a critical yet underexplored challenge, our benchmark includes a dedicated component on hallucination, highlighting how leading models, such as GPT-4o, often hallucinate the existence of a visual cue and erroneously perform the editing task. Our code and benchmark are publicly released at https://github.com/SaraGhazanfari/SpotEdit.

图像编辑幻觉检测生成模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。