构建新基准,评估图像编辑模型的认知与创意能力
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
- 按认知三阶段设计任务:感知、理解、想象
- 涵盖1220个测试用例,揭示当前模型在知识推理上的短板
- 适合研究图像生成、认知建模的学者和开发者
近期图像编辑模型展现出高级智能能力,支持基于认知与创造力的编辑。然而现有基准评估范围过窄,难以全面检验这些能力。为此,我们提出WiseEdit,一个知识密集型基准,用于全面评估认知与创造力驱动的图像编辑。借鉴人类认知创作过程,WiseEdit将图像编辑分解为三个连续步骤:意识(Awareness)、解释(Interpretation)和想象(Imagination),每个步骤对应一项挑战性任务。还包含复杂任务,其中三个步骤均难以完成。该基准融合三类基础知识:陈述性、程序性与元认知知识。最终包含1220个测试用例,客观揭示当前最先进模型在基于知识的认知推理与创造性组合方面的局限性。项目页面、评估代码及各模型生成图像将陆续公开。
原文摘要 · Abstract (English)
Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a scope for evaluation, failing to holistically assess these advanced abilities. To address this, we introduce WiseEdit, a knowledge-intensive benchmark for comprehensive evaluation of cognition- and creativity-informed image editing, featuring deep task depth and broad knowledge breadth. Drawing an analogy to human cognitive creation, WiseEdit decomposes image editing into three cascaded steps, i.e., Awareness, Interpretation, and Imagination, each corresponding to a task that poses a challenge for models to complete at the specific step. It also encompasses complex tasks, where none of the three steps can be finished easily. Furthermore, WiseEdit incorporates three fundamental types of knowledge: Declarative, Procedural, and Metacognitive knowledge. Ultimately, WiseEdit comprises 1,220 test cases, objectively revealing the limitations of SoTA image editing models in knowledge-based cognitive reasoning and creative composition capabilities. The benchmark, evaluation code, and the generated images of each model will be made publicly available soon. Project Page: https://qnancy.github.io/wiseedit_project_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。