arXiv:2602.04347stat.MLcs.LG2026-02中稿 · publication in INF…

用强化学习动态推荐练习题,提升学生个性化学习效果。

A Bandit-Based Approach to Educational Recommender Systems: Contextual Thompson Sampling for Learner Skill Gain Optimization

  • 基于上下文汤普森采样,根据学生历史表现选择最优练习题。
  • 实测推荐练习题带来显著技能提升,优于传统方法。
  • 适合大规模在线教育场景,助教师发现需要支持的学生。

近年来,运筹学、管理科学与分析领域的教学日益转向数字环境,面对大量多样化的学习者,难以提供个性化的练习。本文提出一种生成个性化练习序列的方法:在每一步选择最可能促进学习者掌握目标技能的练习题。该方法利用学习者背景及其过往表现信息进行决策,学习进展以每次练习前后估算技能水平的变化衡量。基于某在线数学辅导平台的数据,结果表明该方法推荐的练习题能带来更大技能提升,并有效适应不同学习者的差异。从教学角度看,该框架实现了规模化个性化练习,凸显具有持续高学习价值的练习题,并帮助教师识别需额外支持的学习者。

原文摘要 · Abstract (English)

In recent years, instructional practices in Operations Research (OR), Management Science (MS), and Analytics have increasingly shifted toward digital environments, where large and diverse groups of learners make it difficult to provide practice that adapts to individual needs. This paper introduces a method that generates personalized sequences of exercises by selecting, at each step, the exercise most likely to advance a learner's understanding of a targeted skill. The method uses information about the learner and their past performance to guide these choices, and learning progress is measured as the change in estimated skill level before and after each exercise. Using data from an online mathematics tutoring platform, we find that the approach recommends exercises associated with greater skill improvement and adapts effectively to differences across learners. From an instructional perspective, the framework enables personalized practice at scale, highlights exercises with consistently strong learning value, and helps instructors identify learners who may benefit from additional support.

推荐系统个性化学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。