arXiv:2409.16202cs.AI2024-09被引 3

用中考题评测大模型,看它真能当学生吗?

CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data

  • 基于26,136道中考真题,覆盖十大学科
  • 包含题型、难度、知识点等详细标注
  • 适合评估教育类大模型实际应用能力

在线教育平台通过数字化基础设施显著推动了教育资源的传播。随着这一转型的深化,大型语言模型(LLMs)提升了平台的智能化水平。然而,现有学术基准对真实产业场景的指导有限,因为教育应用不仅需要答对问题,还需理解上下文与知识脉络。为此,我们提出CJEval,一个基于中国初中升学考试数据的评估基准。CJEval包含26,136个样本,涵盖四类应用级教育任务和十大学科,每个样本均配有题型、难度等级、知识概念及答案解析等详细注释。利用该基准,我们评估了大模型在各类教育任务上的潜力,并通过微调进行性能分析。大量实验与讨论揭示了大模型在教育领域应用的机遇与挑战。

原文摘要 · Abstract (English)

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models (LLMs) has elevated the intelligence levels of these platforms. However, current academic benchmarks provide limited guidance for real-world industry scenarios. This limitation arises because educational applications require more than mere test question responses. To bridge this gap, we introduce CJEval, a benchmark based on Chinese Junior High School Exam Evaluations. CJEval consists of 26,136 samples across four application-level educational tasks covering ten subjects. These samples include not only questions and answers but also detailed annotations such as question types, difficulty levels, knowledge concepts, and answer explanations. By utilizing this benchmark, we assessed LLMs' potential applications and conducted a comprehensive analysis of their performance by fine-tuning on various educational tasks. Extensive experiments and discussions have highlighted the opportunities and challenges of applying LLMs in the field of education.

大模型评测教育AI中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。