基于日本全国学业评估构建90万量级学生作答分布的多模态评测集
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

- 从真实中学考试中提取科学、数学、国语题,保留原始排版与图表
- 包含约90万考生的作答分布数据,支持人机表现直接对比
- 适用于教育类多模态模型评估与可解释AI研究
真实学校考试为评估多模态大语言模型提供了高有效性测试平台,但基于日本中小学评估的基准数据集仍十分稀缺。我们构建了一个基于日本全国学业能力评估的多模态数据集,包含官方发布的中学阶段科学、数学和国语试题。不同于依赖合成或人工筛选数据的现有基准,本数据集保留了真实考试的版式、图表及日本教育文本,并整合了全国范围的考生作答分布(N ≈ 900,000)。这些特性使人类与模型在统一框架下的表现可直接比较。我们使用精确匹配准确率和字符级F1对近期多模态大模型进行评测,发现不同学科间性能差异显著,且对视觉推理需求敏感。通过人工评估与大模型评分器分析进一步验证了自动评分的可靠性。该数据集为多模态教育推理提供了可复现、以人类为基线的评测基准,支持未来在真实评估场景下的评价、反馈生成与可解释AI研究。数据集已公开:https://github.com/KyosukeTakami/gakucho-benchmark
原文摘要 · Abstract (English)
Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japanese K-12 assessments remain scarce. We present a multimodal dataset constructed from Japan's National Assessment of Academic Ability, comprising officially released middle-school items in Science, Mathematics, and Japanese Language. Unlike existing benchmarks based on synthetic or curated data, our dataset preserves real exam layouts, diagrams, and Japanese educational text, together with nationwide aggregated student response distributions (N $\approx$ 900{,}000). These features enable direct comparison between human and model performance under a unified evaluation framework. We benchmark recent multimodal LLMs using exact-match accuracy and character-level F1 for open-ended responses, observing substantial variation across subjects and strong sensitivity to visual reasoning demands. Human evaluation and LLM-as-judge analyses further assess the reliability of automatic scoring. Our dataset establishes a reproducible, human-grounded benchmark for multimodal educational reasoning and supports future research on evaluation, feedback generation, and explainable AI in authentic assessment contexts. Our dataset is available at: https://github.com/KyosukeTakami/gakucho-benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。