arXiv:2506.01676cs.AI2025-06被引 4

构建首个中文K12多模态推理基准,评估模型解题过程与答案。

K12Vista: Exploring the Boundaries of MLLMs in K-12 Education

  • 构建3.3万道题的K12Vista多模态基准,覆盖五大学科三类题型
  • 创建80万条步骤级标注的推理过程数据集,支持细粒度评估
  • 提出首个人工标注的推理过程评估基准,适合教育AI研究者使用

多模态大语言模型在视觉任务中展现出强大推理能力,但在K12教育场景中的能力仍系统性地被低估。以往研究存在学科覆盖窄、数据规模不足、题型单一及仅关注答案的评价方式等局限,难以全面探索模型潜力。为此,我们提出K12Vista,迄今最全面的中文K12学科知识理解与推理多模态基准,涵盖从小学到高中五个核心学科的33,000道题目,包含三种题型。同时,我们关注模型推理过程的正确性,精心收集模型推理错误,并通过自动化数据管道构建K12-PEM-800K,该数据集是目前最大的过程评估数据集,提供详尽的分步判断标注。进一步开发了融合过程与答案一致性的评估模型K12-PEM。此外,引入首个高质量、人工标注的推理过程评估基准K12-PEBench。大量实验表明,当前多模态大语言模型在K12Vista中存在显著缺陷,为更强大模型的开发提供了关键洞见。相关资源已开源:https://github.com/lichongod/K12Vista。

原文摘要 · Abstract (English)

Multimodal large language models have demonstrated remarkable reasoning capabilities in various visual tasks. However, their abilities in K12 scenarios are still systematically underexplored. Previous studies suffer from various limitations including narrow subject coverage, insufficient data scale, lack of diversity in question types, and naive answer-centric evaluation method, resulting in insufficient exploration of model capabilities. To address these gaps, we propose K12Vista, the most comprehensive multimodal benchmark for Chinese K12 subject knowledge understanding and reasoning to date, featuring 33,000 questions across five core subjects from primary to high school and three question types. Moreover, beyond the final outcome, we are also concerned with the correctness of MLLMs' reasoning processes. For this purpose, we meticulously compiles errors from MLLMs' reasoning processes and leverage an automated data pipeline to construct K12-PEM-800K, the largest process evaluation dataset offering detailed step-by-step judgement annotations for MLLMs' reasoning. Subsequently, we developed K12-PEM, an advanced process evaluation model that integrates an overall assessment of both the reasoning process and answer correctness. Moreover, we also introduce K12-PEBench, the first high-quality, human-annotated benchmark specifically designed for evaluating abilities of reasoning process evaluation.Extensive experiments reveal that current MLLMs exhibit significant flaws when reasoning within K12Vista, providing critical insights for the development of more capable MLLMs.We open our resources at https://github.com/lichongod/K12Vista.

多模态模型教育AI推理评估中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。