arXiv:2508.00270cs.LG2025-08被引 1

用百万学生数据优化在线辅导反馈,提升学习效果。

Learning to Optimize Feedback for One Million Students: Insights from Multi-Armed and Contextual Bandits in Large-Scale Online Tutoring

  • 基于多臂赌博机框架,自动选择最优提示策略。
  • 在16.6万次练习中显著提升学生第二次作答正确率。
  • 虽个性化反馈有潜力,但对多数情况提升有限。

我们构建了一个在线辅导系统,通过分析一百万学生的答题数据,学习在学生答错后提供最有效的反馈。采用多臂赌博机(MAB)框架与离线策略评估,共测试43,000种辅助行为(如不同提示),识别出针对不同学习目标(如答对率、任务完成度)的权衡关系。设计算法为每道题动态选择最优训练目标,以提升学生即时重试成功率和整体练习表现。在166,000次练习会话中验证了该策略的有效性,显著改善学生表现。进一步探索上下文赌博机(CB)是否可通过学生特征(如能力估计、反应时间)实现个性化反馈。利用因果推断分析发现:部分题目在特定学生群体中存在异质性影响,但效应量普遍较小,不足以使CB策略显著优于已优化的通用MAB策略。研究揭示大规模数据驱动系统部署的洞察,对未来优化具有指导意义。目前该系统每日支持数千名学生学习。

原文摘要 · Abstract (English)

We present an online tutoring system that learns to provide effective feedback to students after they answer questions incorrectly. Using data from one million students, the system learns which assistance action (e.g., one of multiple hints) to provide for each question to optimize student learning. Employing the multi-armed bandit (MAB) framework and offline policy evaluation, we assess 43,000 assistance actions, and identify trade-offs between assistance policies optimized for different student outcomes (e.g., response correctness, session completion). We design an algorithm that for each question decides on a suitable policy training objective to enhance students' immediate second attempt success and overall practice session performance. We evaluate the resulting MAB policies in 166,000 practice sessions, verifying significant improvements in student outcomes. While MAB policies optimize feedback for the overall student population, we further investigate whether contextual bandit (CB) policies can enhance outcomes by personalizing feedback based on individual student features (e.g., ability estimates, response times). Using causal inference, we examine (i) how effects of assistance actions vary across students and (ii) whether CB policies, which leverage such effect heterogeneity, outperform MAB policies. While our analysis reveals that some actions for some questions exhibit effect heterogeneity, effect sizes may often be too small for CB policies to provide significant improvements beyond what well-optimized MAB policies that deliver the same action to all students already achieve. We discuss insights gained from deploying data-driven systems at scale and implications for future refinements. Today, the teaching policies optimized by our system support thousands of students daily.

在线教育强化学习个性化推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。