arXiv:2504.13143cs.CVcs.AI2025-04被引 28

构建可调控复杂度的图像编辑评测基准,揭示模型在复杂指令下的表现瓶颈。

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

  • 用GPT-4o生成原子编辑任务并组合成复杂指令,形成链式编辑流程。
  • 开源模型在复杂指令下表现远差于闭源模型,且越复杂差距越大。
  • 复杂指令易导致保留原图关键元素和美感能力下降,合成数据加剧此问题。

我们提出Complex-Edit,一个系统性评估指令驱动图像编辑模型在不同复杂度指令下表现的综合基准。通过GPT-4o大规模自动收集多样编辑指令,采用结构化的「编辑链」流程:先生成独立的原子编辑任务,再整合为连贯的复杂指令。同时引入多维度评估指标与基于视觉语言模型(VLM)的自动化评估流水线,支持大规模测试。基准发现:1)开源模型相比闭源模型显著落后,且复杂度越高差距越大;2)复杂指令主要损害模型保留输入图像关键元素和整体美感的能力;3)将复杂指令分解为分步执行的原子操作会大幅降低多指标性能;4)简单的Best-of-N选择策略能提升直接编辑及分步方法的效果;5)观察到“合成数据诅咒”现象:若训练中使用合成数据,随着指令复杂度增加,生成图像越来越不自然——这一现象也出现在最新GPT-4o输出中。

原文摘要 · Abstract (English)

We introduce $\texttt{Complex-Edit}$, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to automatically collect a diverse set of editing instructions at scale. Our approach follows a well-structured ``Chain-of-Edit'' pipeline: we first generate individual atomic editing tasks independently and then integrate them to form cohesive, complex instructions. Additionally, we introduce a suite of metrics to assess various aspects of editing performance, along with a VLM-based auto-evaluation pipeline that supports large-scale assessments. Our benchmark yields several notable insights: 1) Open-source models significantly underperform relative to proprietary, closed-source models, with the performance gap widening as instruction complexity increases; 2) Increased instructional complexity primarily impairs the models' ability to retain key elements from the input images and to preserve the overall aesthetic quality; 3) Decomposing a complex instruction into a sequence of atomic steps, executed in a step-by-step manner, substantially degrades performance across multiple metrics; 4) A straightforward Best-of-N selection strategy improves results for both direct editing and the step-by-step sequential approach; and 5) We observe a ``curse of synthetic data'': when synthetic data is involved in model training, the edited images from such models tend to appear increasingly synthetic as the complexity of the editing instructions rises -- a phenomenon that intriguingly also manifests in the latest GPT-4o outputs.

图像编辑指令评估合成数据GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。