构建真实图像编辑评估基准,全面衡量模型在多图、实用与高阶推理任务中的表现。
CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing

- 设计三类子集:通用、实用与智能评测,覆盖多图编辑与复杂指令场景。
- 主流模型在新基准上性能差异更显著,验证了其区分能力。
- 结果与人类偏好高度一致,适合指导模型优化与实际部署。
随着图像编辑模型快速发展并广泛应用于各领域,亟需将其能力直接部署到真实场景中。然而现有基准仍局限于单一图像任务,覆盖维度有限,难以有效区分不同模型在复杂多图编辑、高难度推理指令及实际部署环境下的表现。为此,我们提出 CPI-Bench——一个综合性、实用性与智能性兼具的图像编辑评估基准。该基准包含三个核心子集:CPI-General-Bench 覆盖多样编辑任务并引入多图评估;CPI-Practical-Bench 关注高频真实用户场景;CPI-Intelligent-Bench 专注于高要求推理型编辑能力评测。基于 CPI-Bench 的主流模型评估结果显示,该基准显著提升了模型间的性能区分度,可全面可靠地量化通用编辑、实际部署与高级推理能力之间的差距,为未来模型优化提供重要参考。关键的是,排名分析表明 CPI-Bench 与 Arena Image Edit Leaderboard 一致性最高,体现出更强的人类偏好对齐,可作为公开人类评价的有效代理。
原文摘要 · Abstract (English)
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and introduces multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings, serving as an effective proxy for public human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。