arXiv:2508.06851cs.AIcs.CY2025-08被引 3

构建多学科测评基准,评估模型在真实考试中的综合能力。

MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

  • 基于14.1万道题构建六层知识体系,覆盖六个学科。
  • 发现现有模型在跨年、跨情境等场景下表现显著下降。
  • 适合关注AI教育应用与模型泛化能力的研究者。

多模态大语言模型(MLLMs)通过融合语言与视觉信息解决复杂问题,对推动通用人工智能(AGI)发展至关重要。然而,当前评估基准普遍存在规模有限、覆盖范围窄、知识结构松散等问题,仅能提供静态且同质化的评测。为此,我们提出MDK12-Bench,一个大规模多学科评估基准,源自涵盖六个学科的真实中小学考试,包含14.1万实例和6,225个知识点,采用六层分类体系组织。该基准覆盖五种题型,并标注难度与年份信息,支持从四个维度全面评估模型表现:1)难度层级,2)时间跨度(跨年)变化,3)上下文迁移,4)基于知识的推理。我们设计了一种新型动态评估框架,引入未见过的视觉、文本及题型变化,以考验模型泛化能力,同时通过减少数据污染提升评估客观性与长期可用性。此外,我们评估了知识点参考增强生成(KP-RAG)在解题中的作用。关键发现揭示了当前MLLMs在多个方面存在局限,为提升模型鲁棒性、可解释性及辅助教育提供了指导。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K-12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in problem-solving. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model robustness, interpretability, and AI-assisted education.

多模态测评基准教育AI知识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。