arXiv:2606.01592cs.CYcs.CL2026-06

LLM生成的英语语法题效果如何?实测发现题型影响学习负担。

Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises

  • 对比三种题型:选择题认知负荷最低,填空最难,拖拽最耗时
  • 学生正确率随CEFR等级提升而下降,反应时间变长,验证框架有效性
  • 适合教育AI开发者优化题目顺序,从识别到输出渐进训练

本研究评估大型语言模型生成的英语作为外语(EFL)学习内容的教育可行性。基于日本初中生在语法练习应用中的日志数据,分析不同题型对学习表现的影响,并检验理论本地化CEFR难度层级是否准确预测实际任务难度。结果表明存在明确的表现层级:选择题认知负荷最低,填空任务对主动回忆构成最大障碍,拖拽题导致最严重的耗时。此外,学习者数据验证了CEFR-J语法框架,显示随着熟练度提升,正确率稳步下降,反应时间增加。研究证明LLM可有效生成学习内容,同时强调开发者需策略性编排题型,帮助学习者从被动识别过渡到主动语言产出。

原文摘要 · Abstract (English)

This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we analysed how different question modalities impact student performance and whether theoretical localised CEFR difficulty tiers accurately predict empirical task difficulty. Results reveal a clear performance hierarchy: multiple-choice questions carried the lowest cognitive load, cloze tasks posed the greatest barrier to active recall, and drag-and-drop exercises incurred the heaviest time penalties. Furthermore, learner data validated the CEFR-J grammar framework, showing a steady decline in accuracy and increased response times as proficiency levels advanced. These findings demonstrate that LLMs can successfully generate learning content, while highlighting the need for developers to strategically sequence question modalities to transition learners from passive recognition to active linguistic production.

语言学习LLM应用教育评估题型设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。