构建艺术史测评新基准,揭示大模型在真实教育场景中的认知短板。
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

- 基于中学与AP艺术史真题,设计多格式混合评估体系。
- 多选题准确率接近满分,但开放填空和错误识别骤降至23.9%与6.2%。
- 强调需结合理由说明,才能真实反映模型的艺术史理解能力。
大语言模型在通用评测中表现接近天花板,但难以揭示其在特定学科中的真实表现。现有艺术类评测多依赖合成问题,且缺乏题目层面的分析。本文提出EduArt,一个面向艺术史知识与视觉推理的教育级多模态评测基准。该基准包含871道来自意大利中学课程与美国大学先修艺术史考试的人工命题,涵盖双语、七种题型,从单选到文本填空、错误识别等。对六家厂商的十二个模型在仅答答案与需提供理由两种条件下进行评估,并通过经典测验理论与逻辑回归分析格式、语言、图像存在性及模型类型的影响。结果表明,该基准具有优良心理测量特性(平均区分度0.514,82.3%为优质区分题),而多选题准确率普遍接近饱和;然而在开放作答与错误识别任务中,部分模型得分暴跌至23.9%与6.2%。要求理由时,各模型表现呈现家族特异性下降趋势。这些差异说明艺术史知识与应用能力是独立能力,单一格式评测会高估模型实际水平。建立全面的能力画像,是负责任使用多模态大模型于艺术史研究的前提,因该领域更注重内容生成与重构,而非选项选择。
原文摘要 · Abstract (English)
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focused evaluations rely on synthetic questions and rarely report item-level properties. This paper introduces EduArt, an educational-level benchmark for art-historical knowledge and visual reasoning in multimodal LLMs. EduArt comprises 871 human-authored questions from Italian secondary-school exercises and US Advanced Placement Art History exams, spanning two languages and seven formats from multiple choice to in-text word placement and error identification. Twelve models from six provider families were evaluated under a default answer-only condition and a motivation condition requiring written justification, and characterized using Classical Test Theory and a logistic regression isolating the effects of format, language, image presence, and model. The benchmark showed strong psychometric properties (mean discrimination 0.514, 82.3 percent good discriminators), while multiple-choice accuracy saturated near ceiling for six models, showing recognition formats alone cannot distinguish frontier models. Format was a strong independent predictor of accuracy: models exceeding 94 percent on multiple choice fell to 23.9 percent on open completion (Claude Opus 4.6) and 6.2 percent on error identification (Claude Sonnet 4.6). The motivation condition changed accuracy in a predominantly negative, family-dependent direction. These dissociations indicate that art-historical knowledge and the ability to deploy it are distinct capabilities, and that single-format benchmarks overestimate what models can reliably do. Mapping this capability profile is a precondition for responsible use of multimodal LLMs in art-historical scholarship, where tasks demand producing and manipulating content rather than selecting from fixed options.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。