构建中文教育大模型评估基准,覆盖知识与素养双重维度。
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- 分知识与素养两大维度,细分为61个学科
- 含24.6万高质量题库,涵盖11种考题类型
- 揭示主流模型在素养题上仍落后人类近30%
随着大语言模型(LLMs)的快速发展,其在教育领域的应用日益广泛。然而,现有模型与评估基准多聚焦于知识维度,忽视了真实教育场景中至关重要的素养能力评估。同时,多数基准局限于单一学科或题型,缺乏多样性,尤其在中文语境下问题突出。为此,我们提出OmniEduBench,一个全面的中文教育评估基准。该基准包含24.602K高质量问答对,按知识与素养两大维度划分,分别含18.121K和6.481K条目。每个维度进一步细分为6个子类别,覆盖总计61个学科(知识类41个,素养类20个)。数据涵盖11种常见考试题型,为全面评估大模型教育能力提供基础。对11个主流开源与闭源模型的实验表明:在知识维度,仅Gemini-2.5 Pro准确率超60%;在素养维度,表现最优的QWQ模型仍比人类低近30%。结果凸显了当前模型在教育应用中的巨大提升空间。
原文摘要 · Abstract (English)
With the rapid development of large language models (LLMs), various LLM-based works have been widely applied in educational fields. However, most existing LLMs and their benchmarks focus primarily on the knowledge dimension, largely neglecting the evaluation of cultivation capabilities that are essential for real-world educational scenarios. Additionally, current benchmarks are often limited to a single subject or question type, lacking sufficient diversity. This issue is particularly prominent within the Chinese context. To address this gap, we introduce OmniEduBench, a comprehensive Chinese educational benchmark. OmniEduBench consists of 24.602K high-quality question-answer pairs. The data is meticulously divided into two core dimensions: the knowledge dimension and the cultivation dimension, which contain 18.121K and 6.481K entries, respectively. Each dimension is further subdivided into 6 fine-grained categories, covering a total of 61 different subjects (41 in the knowledge and 20 in the cultivation). Furthermore, the dataset features a rich variety of question formats, including 11 common exam question types, providing a solid foundation for comprehensively evaluating LLMs' capabilities in education. Extensive experiments on 11 mainstream open-source and closed-source LLMs reveal a clear performance gap. In the knowledge dimension, only Gemini-2.5 Pro surpassed 60\% accuracy, while in the cultivation dimension, the best-performing model, QWQ, still trailed human intelligence by nearly 30\%. These results highlight the substantial room for improvement and underscore the challenges of applying LLMs in education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。