提出评估AI导师教学能力的统一框架,解决评测标准碎片化问题。
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
- 构建涵盖8个维度的教学能力评估体系,基于学习科学原理
- 发布包含192轮对话、1596条回复的MRBench基准数据集
- 验证主流大模型在教学中的实际表现,区分适合作答与教学的模型
本文研究当前最先进大语言模型(LLMs)是否具备作为AI导师的有效性及必要的教学能力,以支持教育对话。以往评估多依赖主观方法和零散基准。为此,我们提出一个统一的评估分类法,包含8个基于学习科学原则的教学维度,用于评估LLM驱动的AI导师在数学领域针对学生错误或困惑时的教学生态价值。我们发布了MRBench——一个新基准,包含7个先进LLM及人类导师的192轮对话、1596条回应,并提供8个维度的标注。我们评估了Prometheus2和Llama-3.1-8B等模型作为评估者的可靠性,分析各导师的教学能力,揭示哪些模型适合做导师,哪些更适合问答系统。我们认为该分类法、基准与人工标注将推动AI导师评估标准化,助力其持续发展。
原文摘要 · Abstract (English)
In this paper, we investigate whether current state-of-the-art large language models (LLMs) are effective as AI tutors and whether they demonstrate pedagogical abilities necessary for good AI tutoring in educational dialogues. Previous efforts towards evaluation have been limited to subjective protocols and benchmarks. To bridge this gap, we propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles, which is designed to assess the pedagogical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain. We release MRBench - a new evaluation benchmark containing 192 conversations and 1,596 responses from seven state-of-the-art LLM-based and human tutors, providing gold annotations for eight pedagogical dimensions. We assess reliability of the popular Prometheus2 and Llama-3.1-8B LLMs as evaluators and analyze each tutor's pedagogical abilities, highlighting which LLMs are good tutors and which ones are more suitable as question-answering systems. We believe that the presented taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors' development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。