arXiv:2605.26781cs.AIcs.MM2026-05

测试大模型在真实考试中的表现,发现其能力远未达标。

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

论文配图:LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
图 1 · 摘自论文原文
  • 构建动态真实考试数据集,自动更新避免数据泄露。
  • 模型在真实考试环境下得分下降26分,暴露实际应用短板。
  • 适合教育科技、AI评估研究者参考,关注真实场景性能。

先进大型多模态模型(LMMs)在中小学推理任务中表现出色,具备作为智能导师的巨大潜力。但要实现这一目标,需确保模型能在真实考试环境中有效应对。现有大多数基准存在静态、数据污染、模态和学科受限等问题。为此,我们提出LiveK12Bench——一个动态、全面、跨学科的基准,用于评估LMM在真实考试场景下的推理能力。该基准包含2000+经验证的问题,涵盖数学、物理、化学、生物,源自最新真实考试试卷,并支持持续更新。核心创新包括:1)自动化流水线持续接入并解析最新考卷,防止数据泄露;2)提出新型「模拟考试」评估方案,评估模型自主完成整套考试的能力,要求推理路径准确且高效。对12个LMM的实验表明,在真实考试约束下,先进模型性能显著下降:GPT-5得分从79降至53(满分100)。研究揭示了对复杂视觉布局敏感等关键缺陷,暴露出理想推理能力与真实教育适应性之间的差距。代码与数据集已公开。

原文摘要 · Abstract (English)

Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations effectively, yet most existing benchmarks fail to capture the complexity of authentic testing environments. Specifically, most datasets are static, prone to data contamination, and are often confined to restricted modalities, disciplines, and evaluation criteria. To address these issues, we introduce LiveK12Bench, a dynamic, holistic, multi-disciplinary benchmark designed to evaluate the reasoning abilities of LMMs in realistic examination scenarios. LiveK12Bench comprises 2K+ verified questions spanning Mathematics, Physics, Chemistry, and Biology, sourced from the latest real-world exam papers and designed to grow over time. Our framework features several core innovations: 1) featuring an automated pipeline that continuously ingests and parses the latest examination papers to mitigate data leakage; and 2) proposing a novel `Mock Exam' evaluation scheme, which assesses the ability to complete end-to-end exams autonomously with accurate and efficient reasoning paths. Extensive experiments on 12 LMMs reveal that advanced models suffer substantial performance degradation under exam-realistic constraints: GPT-5's score drops from 79 to 53 (out of 100) when process rigor and efficiency are jointly evaluated. Our findings expose critical vulnerabilities, such as sensitivity to complex visual layouts, highlighting the gap between idealized reasoning capabilities and true educational readiness. Both code and dataset are publicly available.

多模态模型教育AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。