arXiv:2607.14591cs.CL2026-07

测试2万篇留学生作文的AI反馈效果,发现专家评分和学生感受严重不符。

How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

论文配图:How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts
图 1 · 摘自论文原文
  • 用教师评分与学生反馈双视角评估AI生成的英语作文修改意见
  • 2000多名学生参与,2万份作文反馈显示专家与学生评价差异大
  • 强调教育类AI需从学习者角度设计评估体系

本研究聚焦英语作为外语(EFL)写作中的书面纠正性反馈(WCF)。大型语言模型(LLMs)可大规模提供WCF,但其与教学最佳实践的对齐仍是挑战。即使生成的反馈在事实性或相关性上达标,仍可能不适用于学习场景,凸显基于学习者视角的外在评估的重要性。我们在一所大学的EFL课程中部署了WCF系统,收集了近2000名学生的超过20,000份作文草稿。从两个角度评估生成的WCF:一是由经验丰富的英语教师使用量表进行内在评估,二是通过学生反馈和参与度指标进行外在评估。结果表明,教师专家评分与学生反馈之间存在较低一致性。这说明仅依赖传统专家评估无法全面反映AI反馈在学习情境中的可用性或帮助程度,强调在语言教育领域构建以学习者为中心的评估框架至关重要。

原文摘要 · Abstract (English)

This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.

AI教育语言学习反馈评估大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。