评测图像编辑模型的推理能力,发现现有模型在知识理解上仍有巨大短板。
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
- 按事实、概念、程序三类知识设计22项任务
- 在10个顶尖模型上测试,发现推理能力普遍不足
- 适合研究智能图像编辑与认知评估的学者
近期多模态生成模型在基于指令的图像编辑方面取得显著进展,但其在知识驱动的推理型编辑任务上的表现仍待深入探索。本文提出KRIS-Bench(基于知识推理的图像编辑系统基准),一个基于认知理论设计的诊断性评测基准。该基准依据教育学理论,将编辑任务分为三类基础性知识:事实型、概念型和程序型,据此构建了覆盖7个推理维度的22个代表性任务,并发布1,267个高质量标注的编辑实例。为支持细粒度评估,我们提出一套完整评测协议,包含一种新的知识合理性度量(Knowledge Plausibility),结合知识提示并经人工验证校准。对10个先进模型的实证结果表明,各模型在推理能力上存在显著差距,凸显了面向知识的核心评测基准对推动智能图像编辑系统发展的必要性。
原文摘要 · Abstract (English)
Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。