arXiv:2506.17006cs.CLcs.CY2025-06中稿 · EC-TEL '25被引 12

LLM生成的反馈能提升学习效果,但只对主动使用的人有效。

LLM-Generated Feedback Supports Learning If Learners Choose to Use It

  • 让学习者自主选择是否使用GPT-3.5生成的解释性反馈。
  • 两门课程中使用LLM反馈者成绩提升显著,效应量0.28~0.33。
  • 适合想在开放任务中低成本提升学习的教育系统参考。

大型语言模型(LLMs)越来越多地用于生成反馈,但其对学习的影响仍缺乏深入研究,尤其与传统反馈方式相比。本研究考察了在七场基于情境的导师培训课程中,按需生成的解释性反馈对学习的影响。分析了来自885名学习者的2600多次课程完成数据,比较了三组表现:使用GPT-3.5-turbo生成反馈者、拒绝使用者及无访问权限者。所有组均获得非LLM纠正反馈。为缓解高绩效者更倾向使用LLM反馈的选择偏差,采用倾向评分法。预测高使用倾向的学习者在后测中表现显著优于低倾向者。经调整后,七门课程中有两门显示统计学上显著的学习增益,标准化效应量分别为0.28和0.33。这表明LLM反馈的效果取决于学习者寻求支持的意愿。重要的是,使用LLM反馈未显著增加完成时间,且学习者普遍评价其有帮助。研究结果凸显了LLM反馈在开放式任务中作为低成本、可扩展学习增强工具的潜力,尤其适用于已有反馈系统的现有平台。本工作公开了数据集、提示词与评分标准以支持可复现性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to generate feedback, yet their impact on learning remains underexplored, especially compared to existing feedback methods. This study investigates how on-demand LLM-generated explanatory feedback influences learning in seven scenario-based tutor training lessons. Analyzing over 2,600 lesson completions from 885 tutor learners, we compare posttest performance among learners across three groups: learners who received feedback generated by gpt-3.5-turbo, those who declined it, and those without access. All groups received non-LLM corrective feedback. To address potential selection bias-where higher-performing learners may be more inclined to use LLM feedback-we applied propensity scoring. Learners with a higher predicted likelihood of engaging with LLM feedback scored significantly higher at posttest than those with lower propensity. After adjusting for this effect, two out of seven lessons showed statistically significant learning benefits from LLM feedback with standardized effect sizes of 0.28 and 0.33. These moderate effects suggest that the effectiveness of LLM feedback depends on the learners' tendency to seek support. Importantly, LLM feedback did not significantly increase completion time, and learners overwhelmingly rated it as helpful. These findings highlight LLM feedback's potential as a low-cost and scalable way to improve learning on open-ended tasks, particularly in existing systems already providing feedback without LLMs. This work contributes open datasets, LLM prompts, and rubrics to support reproducibility.

LLM反馈学习增强教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。