构建复杂指令图像编辑基准,提升模型精细操作能力评估
CompBench: Benchmarking Complex Instruction-guided Image Editing
- 采用多模态大模型与人工协作框架构建数据集
- 提出指令解耦策略,拆分编辑意图为四大维度
- 揭示现有模型在空间与上下文推理上的根本缺陷
真实应用场景对复杂场景操控的需求日益增长,但现有指令引导图像编辑基准往往简化任务复杂度,缺乏全面且细粒度的指令。为此,我们提出CompBench,一个面向复杂指令引导图像编辑的大规模基准。该基准包含需精细指令遵循、空间与上下文推理的挑战性编辑场景,可全面评估图像编辑模型的精确操控能力。为构建CompBench,我们设计了多模态大模型与人类协作的框架及定制化任务流程。此外,我们提出指令解耦策略,将编辑意图拆分为位置、外观、动态和对象四个关键维度,确保指令与复杂编辑需求更紧密对齐。大量评估表明,CompBench揭示了当前图像编辑模型的根本局限,并为下一代指令引导图像编辑系统的发展提供关键洞见。项目页面见 https://comp-bench.github.io/。
原文摘要 · Abstract (English)
While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap, we introduce CompBench, a large-scale benchmark specifically designed for complex instruction-guided image editing. CompBench features challenging editing scenarios that incorporate fine-grained instruction following, spatial and contextual reasoning, thereby enabling comprehensive evaluation of image editing models' precise manipulation capabilities. To construct CompBench, we propose an MLLM-human collaborative framework with tailored task pipelines. Furthermore, we propose an instruction decoupling strategy that disentangles editing intents into four key dimensions: location, appearance, dynamics, and objects, ensuring closer alignment between instructions and complex editing requirements. Extensive evaluations reveal that CompBench exposes fundamental limitations of current image editing models and provides critical insights for the development of next-generation instruction-guided image editing systems. Our project page is available at https://comp-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。