arXiv:2506.11604cs.AIcs.CL2025-06中稿 · ed被引 5

测试AI看图理解德国中学生知识,发现顶尖模型准确率不足45%。

VLM@school -- Evaluation of AI image understanding on German middle school knowledge

  • 基于德国中学真实课程设计图文推理题
  • 13个主流模型平均准确率低于45%,数学音乐尤差
  • 适合评估非英语语境下AI的多模态理解能力

本文提出一个新基准数据集,用于评估视觉语言模型(VLMs)在结合视觉推理与特定学科背景知识的德语任务中的表现。与依赖人工难题或去情境化问题的常用英文基准不同,该数据集源自德国九大学科领域的真实中学课程,涵盖数学、历史、生物、宗教等。数据集包含超过2000道开放性问题,基于486张图像,要求模型将视觉解读与事实推理相结合,而非依赖表面文本线索。我们在多个维度上评估了13个先进的开源VLMs,包括领域特定准确率和对抗性问题下的表现。结果显示,即使最强模型整体准确率也低于45%,尤其在音乐、数学及对抗性设置下表现不佳。结果表明,现有模型在主流基准上的成功与其在真实多模态理解任务中的表现存在显著差距。我们得出结论:中学水平任务为压力测试VLMs提供了有意义且未被充分开发的路径,尤其是在非英语语境下。该数据集与评估协议可作为严格测试未来AI系统视觉与语言推理能力的基准。

原文摘要 · Abstract (English)

This paper introduces a novel benchmark dataset designed to evaluate the capabilities of Vision Language Models (VLMs) on tasks that combine visual reasoning with subject-specific background knowledge in the German language. In contrast to widely used English-language benchmarks that often rely on artificially difficult or decontextualized problems, this dataset draws from real middle school curricula across nine domains including mathematics, history, biology, and religion. The benchmark includes over 2,000 open-ended questions grounded in 486 images, ensuring that models must integrate visual interpretation with factual reasoning rather than rely on superficial textual cues. We evaluate thirteen state-of-the-art open-weight VLMs across multiple dimensions, including domain-specific accuracy and performance on adversarial crafted questions. Our findings reveal that even the strongest models achieve less than 45% overall accuracy, with particularly poor performance in music, mathematics, and adversarial settings. Furthermore, the results indicate significant discrepancies between success on popular benchmarks and real-world multimodal understanding. We conclude that middle school-level tasks offer a meaningful and underutilized avenue for stress-testing VLMs, especially in non-English contexts. The dataset and evaluation protocol serve as a rigorous testbed to better understand and improve the visual and linguistic reasoning capabilities of future AI systems.

视觉语言模型多模态理解教育评估德语AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。