arXiv:2506.12830cs.CV2025-06被引 24

新基准评估复杂指令图像编辑,提升模型理解链式操作能力

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

  • 构建复杂指令编辑基准,包含多步依赖关系的链式任务
  • 提出新一致性评估法,精准检测未修改区域的视觉一致性
  • 基于思维链方法显著提升模型对复杂指令的执行效果

文本驱动的图像编辑在处理单一指令上已取得显著进展,但真实场景中常涉及复杂、多步骤的指令,尤其是操作间存在依赖关系的“链式”指令。现有模型对此类复杂指令表现不佳,且现有基准未能充分评估此类能力,往往忽略多指令与链式依赖结构,且常用的一致性度量存在缺陷。为此,我们提出ComplexBench-Edit,一个系统化评估复杂多指令与链式依赖图像编辑任务的新基准。该基准引入一种新视觉一致性评估方法,通过排除已编辑区域,准确评估未修改区域的保持程度。此外,我们提出一种简单而强大的基于思维链(CoT)的方法,显著增强现有模型对复杂指令的理解与执行能力。大量实验表明,ComplexBench-Edit能有效区分模型性能差异,且我们的CoT方法在处理复杂编辑任务中表现更优。数据与代码已开源。

原文摘要 · Abstract (English)

Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ``chain'' instructions where operations are interdependent. Current models struggle with these intricate directives, and existing benchmarks inadequately evaluate such capabilities. Specifically, they often overlook multi-instruction and chain-instruction complexities, and common consistency metrics are flawed. To address this, we introduce ComplexBench-Edit, a novel benchmark designed to systematically assess model performance on complex, multi-instruction, and chain-dependent image editing tasks. ComplexBench-Edit also features a new vision consistency evaluation method that accurately assesses non-modified regions by excluding edited areas. Furthermore, we propose a simple yet powerful Chain-of-Thought (CoT)-based approach that significantly enhances the ability of existing models to follow complex instructions. Our extensive experiments demonstrate ComplexBench-Edit's efficacy in differentiating model capabilities and highlight the superior performance of our CoT-based method in handling complex edits. The data and code are released at https://github.com/llllly26/ComplexBench-Edit.

图像编辑链式指令基准测试思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。