首个面向教育场景的综合评测数据集,支持多维度模型评估
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
- 构建涵盖9类教育场景的4000+上下文合成数据集
- 提出12维评估指标并经人工标注验证效果
- 小模型在该数据集上表现媲美顶尖大模型
随着大语言模型的持续发展,其在教育场景中的应用仍处于探索阶段且未充分优化。本文填补这一空白,推出首个专为教育场景设计的综合性基准数据集,包含9种主要教育场景和超过4000个不同的教育上下文。为实现全面评估,我们提出了覆盖12个关键维度的多维度评价指标,涵盖教师与学生关注的核心方面。通过人工标注验证了模型生成评价结果的有效性。此外,我们在该数据集上训练了一个相对小规模的模型,并证明其在测试集上的表现可与当前顶尖大模型(如Deepseek V3、Qwen Max)相媲美。本研究为面向教育的大语言模型开发与评估提供了实用基础。代码与数据已开源:https://github.com/ybai-nlp/EduBench。
原文摘要 · Abstract (English)
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation metrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-of-the-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models. Code and data are released at https://github.com/ybai-nlp/EduBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。