评测大模型辅导能力的基准数据集,发现当前顶尖模型表现仍不理想。
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
- 构建1490个高中生和AP课程辅导样本,涵盖三类核心教学任务。
- 16个前沿大模型平均得分不足56%,诊断与个性化指导能力普遍欠缺。
- 首次提供细粒度评分标准,适合研究AI教育助手的团队使用。
随着学生越来越多地将大语言模型(LLMs)作为学习辅助工具,构建能够处理辅导复杂性的模型至关重要:需识别学生核心需求、具备自适应性、提供个性化指导并保证准确性。为此,我们推出了TutorBench——一个用于严格评估LLMs核心辅导能力的数据集与评估基准。该数据集包含由人类专家精心设计的1,490个样本,覆盖高中及AP水平课程内容,聚焦三大常见辅导任务:(i) 针对学生产生困惑生成自适应解释;(ii) 对学生作业提供可操作反馈;(iii) 通过有效提示促进主动学习。为应对辅导的复杂性,每个样本均配有专属评分标准,用于评估模型输出。TutorBench采用基于大模型评分员与样本特定标准的可靠且细粒度自动评估方法。我们在TutorBench上评估了16个前沿大模型,并进行了详细性能分析。结果表明,无一模型得分超过56%,提升空间巨大。所有模型在引导、诊断与支持学生所需的核心辅导技能上表现均不理想,相关评分标准通过率均低于60%。不同模型家族表现出差异性优劣:Claude系列在促进主动学习方面表现最佳,但在其他两类任务中落后。通过发布TutorBench,我们提供了一个全面且尚未饱和的基准,以推动下一代AI导师的发展。
原文摘要 · Abstract (English)
As students increasingly adopt large language models (LLMs) as learning aids, it is crucial to build models that are adept at handling the nuances of tutoring: they need to identify the core needs of students, be adaptive, provide personalized guidance, and be accurate. To this end, we introduce TutorBench, a dataset and evaluation benchmark designed to rigorously evaluate the core tutoring skills of LLMs. The dataset comprises 1,490 samples curated by human experts, focused on high-school and AP-level curricula. The samples are drawn from three common tutoring tasks: (i) generating adaptive explanations tailored to a student's confusion, (ii) providing actionable feedback on a student's work, and (iii) promoting active learning through effective hint generation. To account for the inherent complexity of tutoring, samples are accompanied by sample-specific rubrics which are used to judge model responses during evaluation. TutorBench uses a reliable and fine-grained automatic evaluation method that uses an LLM-judge and the sample-specific rubrics. We evaluate 16 frontier LLMs on TutorBench and present a detailed analysis of their performance and behavior. Our results show that none of the frontier LLMs achieve a score of greater than $56\%$, showing a large room for improvement. We find that LLMs fall short in exhibiting the full range of tutoring skills needed to guide, diagnose, and support students effectively, with all the frontier models achieving less than a $60\%$ pass rate on rubric criteria related to these skills. We also find that different model families exhibit varied strengths and limitations: the Claude models outperform others in supporting active learning, while they lag behind in the other two use cases. By releasing TutorBench, we provide a comprehensive and unsaturated benchmark to guide the development of the next-generation of AI tutors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。