评测大模型导师的辅导能力,发现仍有提升空间。
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
- 基于学习科学设计多维度评估框架,量化辅导表现。
- 最佳模型在错误识别上达到71.81的宏平均F1,指导能力仅58.34。
- 公开全部数据与评测工具,助力教育AI研究。
本共享任务旨在评估基于大语言模型的AI导师在教育对话中的教学能力,重点考察其针对学生错误进行纠正回应的质量。任务包含五个赛道,分别评估错误识别、错误定位、提供指导及反馈可操作性等关键维度,并依据学习科学原则定义优质辅导行为;另设导师身份检测赛道。共吸引全球50余支队伍参与。所有模型均以人工标注的黄金标准进行评估,结果虽具潜力,但仍有显著提升空间:四个教学能力赛道的最佳宏平均F1分数介于58.34(提供指导)至71.81(错误识别)之间,而导师身份识别赛道在9分类任务中最高达96.98。本文综述任务主要发现,分析各团队方法与性能表现,并公开所有资源以支持该领域后续研究。
原文摘要 · Abstract (English)
This shared task has aimed to assess pedagogical abilities of AI tutors powered by large language models (LLMs), focusing on evaluating the quality of tutor responses aimed at student's mistake remediation within educational dialogues. The task consisted of five tracks designed to automatically evaluate the AI tutor's performance across key dimensions of mistake identification, precise location of the mistake, providing guidance, and feedback actionability, grounded in learning science principles that define good and effective tutor responses, as well as the track focusing on detection of the tutor identity. The task attracted over 50 international teams across all tracks. The submitted models were evaluated against gold-standard human annotations, and the results, while promising, show that there is still significant room for improvement in this domain: the best results for the four pedagogical ability assessment tracks range between macro F1 scores of 58.34 (for providing guidance) and 71.81 (for mistake identification) on three-class problems, with the best F1 score in the tutor identification track reaching 96.98 on a 9-class task. In this paper, we overview the main findings of the shared task, discuss the approaches taken by the teams, and analyze their performance. All resources associated with this task are made publicly available to support future research in this critical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。