首个评估图像编辑中学科知识推理的基准,覆盖10大学科
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
- 构建跨10个学科的520个精细样本集,测试模型在专业领域推理能力
- 发现当前模型在隐含知识驱动编辑任务中表现严重不足,差距显著
- 适合研究多模态推理、通用视觉模型与教育/科研场景应用者
统一多模态模型旨在实现理解、推理与生成的协同,但现有图像编辑基准主要局限于自然图像和浅层常识推理,难以评估其在结构化、领域特定约束下的表现。本文提出GRADE,首个评估图像编辑中学科知识与推理能力的基准。GRADE包含跨越10个学术领域的520个精心设计样本,涵盖自然科学到社会科学。为支持严谨评估,我们提出多维度评价协议,联合考察学科推理、视觉一致性与逻辑可读性。对20个先进开源及闭源模型的广泛实验揭示了当前模型在隐含知识密集型编辑场景下存在显著局限,性能差距巨大。除量化指标外,还进行深入分析与消融实验,暴露模型缺陷并识别学科编辑中的制约因素。GRADE指明了未来统一多模态模型发展的关键方向,推动学科导向图像编辑与推理研究。本基准及评估代码已公开。
原文摘要 · Abstract (English)
Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain-specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline-informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi-dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in current models under implicit, knowledge-intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline-informed image editing and reasoning. Our benchmark and evaluation code are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。