测试大模型能否像老师一样纠正英语学习错误并讲清原因
Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?
- 用参数调整和检索增强生成测试模型纠错能力
- 表面纠错不错,但解释缺乏专业术语和教学深度
- 适合教育科技研究者和对AI教学效果存疑的教师
尽管越来越多机构鼓励在课堂使用大语言模型,但对其在语言教学核心任务中的实际表现仍缺乏严谨、系统的评估。本文检验当前最先进的大模型是否能提供语言学习者所需的纠错反馈与教学解释。通过系统调整模型参数,研究其对输出质量、教学清晰度和一致性的影响,并结合检索增强生成查询教学方法数据。评估采用自动指标(GLEU、BERTScore)及人工专家判断,以捕捉纯计算方法难以衡量的语言细微差别、文化敏感性和教学适宜性。结果显示,模型虽具备出色的表层纠错能力,但其解释常缺乏有效的术语体系与领域知识,表明当前对AI辅助语言学习的热情可能超前于对其真实教学能力的理解。
原文摘要 · Abstract (English)
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。