用强化学习主动调整习题顺序,提升学生编程成绩
Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

- 将大模型与强化学习结合,根据互动数据动态选题
- 实验显示成绩提升0.15个标准差,相当于多学6-9个月
- 适合教育科技开发者和个性化学习研究者
生成式人工智能(GenAI)正重塑教育,但现有平台多为被动应答的聊天机器人。我们提出,通过主动引导学习可显著提升其效果。为此,设计了一个新教学平台,将精心设计的GenAI聊天机器人与强化学习算法结合,用于自适应排序练习题。该算法利用学生与聊天机器人的互动信号,动态选择难度适中的题目。在台北市府与美国在台协会合作下,于十所高中开展为期五个月的Python课程实验,随机分配学生至固定题序组与自适应题序组。结果显示,自适应序列使未辅助期末考试成绩提高0.15个标准差(据估算相当于6-9个月学业进展);中介分析表明,提升主要源于参与度增加。本研究提供了大规模实地证据,证明学生-聊天机器人互动可作为主动优化与个性化学习的关键信号。
原文摘要 · Abstract (English)
Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely focus on GenAI chatbot tutors that reactively answer student questions. We hypothesize that the efficacy of GenAI chatbot tutors can be substantially improved by proactively guiding student learning. To test this, we design a novel tutoring platform that tightly integrates a carefully-designed GenAI chatbot with a reinforcement learning algorithm for sequencing practice problems. Critically, this algorithm leverages rich signals from student-chatbot interactions to adaptively select practice problems of an appropriate difficulty level. In partnership with the Taipei City Government and American Institute in Taiwan, we deployed our tutoring platform in conjunction with a five-month course to teach Python to students across ten high schools. We randomized students between a fixed practice problem sequence and our adaptive sequencing algorithm. We find that adaptive sequencing increased unassisted final exam performance by 0.15 standard deviations (equivalent to 6-9 months of schooling by some estimates); mediation analysis suggests that gains were driven by increased engagement. Our work provides large-scale field evidence that student-chatbot interactions provide valuable signals for proactively optimizing and personalizing student learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。