arXiv:2607.10647cs.CL2026-07

用知识蒸馏训练小模型,自动评估AI助教教学质量。

Knowledge Distillation for Automated AI Tutor Evaluation

论文配图:Knowledge Distillation for Automated AI Tutor Evaluation
图 1 · 摘自论文原文
  • 用大模型蒸馏生成标注数据,解决教学评估数据少的问题。
  • 在四项教学能力评测中平均提升22.63个百分点,最佳模型达82.88%。
  • 适合教育AI研发者和评估者,可自动化对比主流AI助教表现。

大型语言模型(LLM)在中小学及高等教育中的快速应用,已超过其教学品质评估方法的发展速度。随着研究社区开始探索自动化评估AI助教的路径,本文提出FATE(FLC AI Tutor Evaluator),一个专用于评估的80亿参数语言模型。该模型遵循BEA 2025共享任务的四大核心评价维度:错误识别、错误定位、指导性与可操作性。由于教学评估任务数据有限,我们采用来自前沿大模型的知识蒸馏技术生成额外监督信号,实现最高达22.63个百分点的性能提升。最后,通过基准测试主流商业模型(包括ChatGPT、Claude、Gemini、DeepSeek)生成的教学响应,验证了FATE的实用性。结果显示,Gemini 2.5 Flash表现最优(82.88%),其次为ChatGPT 5.5 Instant(80.75%)、DeepSeek V4 Flash(80.13%)和Claude Sonnet 4.6(74.00%)。

原文摘要 · Abstract (English)

The rapid integration of Large Language Models (LLMs) into K-12 and higher education has outpaced the development of reliable methods for evaluating their pedagogical quality. As the research community starts to explore the space of automating evaluation of AI tutors, we introduce FATE (FLC AI Tutor Evaluator), a specialized 8B-parameter language model designed to evaluate AI tutors. Aligned with the four core evaluation tracks from the BEA 2025 Shared Task, our model assesses pedagogical ability across Mistake Identification, Mistake Location, Guidance, and Actionability. Because pedagogical evaluation is a specialized task with limited labeled data, we leverage knowledge distillation from a frontier LLM to generate additional supervision, yielding absolute performance gains up to 22.63 percentage points. Finally, we demonstrate FATE's utility as an automated evaluator by benchmarking instructional responses generated by popular commercial models, including ChatGPT, Claude, Gemini, and DeepSeek. On average, we have found that Gemini 2.5 Flash perfomed best (82.88%), then ChatGPT 5.5 Instant (80.75%), followed by DeepSeek V4 Flash (80.13%) and Claude Sonnet 4.6 (74.00%).

AI助教知识蒸馏教育评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。