用学生实际行为数据评估AI导师效果,发现传统方法遗漏关键信息
The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness

- 基于10235份代码提交数据,分析学生是否采纳反馈及采纳是否正确
- 发现两种部署的AI导师在学生参与模式上差异显著,传统评价无法捕捉
- 学生行为信号比教学质量更能预测反馈是否被觉得有用,适合教育技术研究者
当前基于人工智能的辅导系统(AI导师)主要依据反馈内容的教学质量进行评估。然而,仅关注教学品质是不够的,因为忽略了核心问题:学生实际上如何响应收到的反馈?本文主张,应引入以学生交互数据为基础的行为维度来补充教学评估。为此提出一种评估框架,并应用于某本科编程课程中10,235份代码提交及其对应AI导师反馈的数据,测量学生是否采纳反馈以及采纳是否准确。通过该框架对不同学期部署的两个AI导师进行比较,揭示出显著不同的学生参与模式,这些差异在仅靠教学评估时未被察觉。此外,基于行为的信号与学生对反馈帮助性的感知关联度高于教学品质本身,为评估AI导师性能提供了更完整、更具行动意义的图景。
原文摘要 · Abstract (English)
Current Artificial Intelligence (AI)-based tutoring systems (AI tutors) are primarily evaluated based on the pedagogical quality of their feedback messages. While important, pedagogy alone is insufficient because it ignores a critical question: what do students actually do with the feedback they receive? We argue that AI tutor evaluation should be extended with a behavioral dimension grounded in student interaction data, which complements pedagogical assessment. We propose an evaluation framework and apply it to 10,235 code submissions with corresponding AI tutor feedback from an introductory undergraduate programming course to measure whether students act on tutor feedback and whether those actions are applied correctly. Using this framework to compare two deployed AI tutors across different semesters in a large-scale introductory computer science course reveals substantial differences in student engagement patterns that are not captured by pedagogy-only evaluation. Moreover, these engagement-based behavioral signals are more strongly associated with student perception of helpful feedback than pedagogical quality alone, providing a more complete and actionable picture of AI tutor performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。