用大模型实现英语写作语法动态评估,支持大规模个性化反馈。
Large Language Model-Driven Dynamic Assessment of Grammatical Accuracy in English Language Learner Writing
- 构建模块化系统DynaWrite,集成多模型生成动态反馈
- GPT-4o在错误识别准确率上与神经聊天相当,但提示质量更优
- 实测响应快、稳定性强,适合课堂规模化应用
本研究探索大型语言模型(LLM)在扩展动态评估(DA)中的潜力。为此,我们开发了DynaWrite——一个基于微服务架构的语法辅导应用,支持多种LLM生成学习者英语写作的动态反馈。对21个LLM的初步测试表明,GPT-4o和neural chat最具发展潜力。进一步对比发现,两者在识别用户句子语法错误方面表现相近;但GPT-4o在生成清晰、一致且逐步明确的提示上显著优于neural chat。性能测试确认系统具备实时响应与稳定性,其中GPT-4o表现出足够速度与可靠性。结果表明,LLM可有效扩展动态评估,使其在传统师生互动之外实现大规模部署。
原文摘要 · Abstract (English)
This study investigates the potential for Large Language Models (LLMs) to scale-up Dynamic Assessment (DA). To facilitate such an investigation, we first developed DynaWrite-a modular, microservices-based grammatical tutoring application which supports multiple LLMs to generate dynamic feedback to learners of English. Initial testing of 21 LLMs, revealed GPT-4o and neural chat to have the most potential to scale-up DA in the language learning classroom. Further testing of these two candidates found both models performed similarly in their ability to accurately identify grammatical errors in user sentences. However, GPT-4o consistently outperformed neural chat in the quality of its DA by generating clear, consistent, and progressively explicit hints. Real-time responsiveness and system stability were also confirmed through detailed performance testing, with GPT-4o exhibiting sufficient speed and stability. This study shows that LLMs can be used to scale-up dynamic assessment and thus enable dynamic assessment to be delivered to larger groups than possible in traditional teacher-learner settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。