评测AI导师需超越对错,GRADE提出系统方法提升教学评估能力。
GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors

- 构建多维度评估框架,融合思维链与微调技术提升教学判断力。
- Gemma3-27B在8位精度下多任务表现最优,优于多数商用系统。
- 揭示微调会干扰指令遵循,推理模式显著影响碳排放与性能。
评估AI导师回应不仅需事实正确:还需识别错误、定位问题、提供指导并给出可操作下一步。本文提出GRADE,针对学生-导师对话中的教学能力评估开展系统研究。基于BEA 2025 TutorMind设置,评估120种配置,涵盖五种语言模型,采用零样本推理、LoRA微调、合成数据增强、CoT+推理及单任务与多任务范式。Gemma3-12B在单任务中表现最佳,而Gemma3-27B在8位精度下多任务预测更可靠。发现数据增强有助于原数据表现弱的模型,验证虽成本高但增益有限,且CoT+推理更适用于生成合成数据而非直接分类。进一步表明,基于结构化分类目标的LoRA微调会干扰思考模式下的指令遵循行为,导致生成偏离预期格式。碳足迹分析显示模型选择与推理模式显著影响碳排放。总体而言,精心选择的开源LoRA流程可在关键教学维度上媲美甚至超越专有与集成系统,代码与数据已公开于https://github.com/pvbgeek/GRADE。
原文摘要 · Abstract (English)
Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionable next steps. We present GRADE, a systematic study of open-source models for pedagogical ability assessment in student-tutor dialogues. Building on the BEA 2025 TutorMind setting, we evaluate 120 configurations across five language models, zero-shot inference, LoRA fine-tuning, synthetic augmentation, CoT+Reasoning, and single-task versus multitask formulations. Gemma3-12B performs best for single-task evaluation, while Gemma3-27B in 8-bit precision is more reliable for multitask prediction. We find that augmentation helps models that struggle with the original data, verification adds limited gains despite higher cost, and CoT+Reasoning is more useful for synthetic data generation than direct classification. We further show that LoRA fine-tuning on structured classification objectives interferes with instruction-following behavior under thinking mode, redirecting generation away from the required evaluation format. Carbon analysis shows that model choice and reasoning mode substantially affect emissions. Overall, GRADE shows that carefully selected open-source LoRA pipelines can match or surpass proprietary and ensemble-based systems on key pedagogical dimensions, with code and data available at https://github.com/pvbgeek/GRADE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。