arXiv:2606.23015cs.LG2026-06

用历史数据训练自适应教学策略,无需新实验即可提升学习效果。

Counterfactual learning of new adaptive instructional policies using logged data

论文配图:Counterfactual learning of new adaptive instructional policies using logged data
图 1 · 摘自论文原文
  • 基于学生答题数据构建连续能力-难度模型,将教学转化为连续强化学习问题。
  • 在4个真实数据集上,新策略相比原有策略显著提升学习成效。
  • 计算快速,支持实时优化,适合教育系统快速迭代改进。

智能辅导系统(ITS)优化教学策略通常依赖昂贵的在线实验或难以反映真实场景的学生模拟器。本文提出一种离线上下文动作框架,直接从已有的交互日志数据中学习新的自适应教学策略。通过使用Rasch模型将学生与题目互动映射到连续的能力-难度潜空间,将教学过程建模为连续随机多臂老虎机问题。设计了一种新型奖励函数,旨在通过平衡任务挑战性与学生成功率来优化“心流”体验。方法包含针对每一轮行为策略的估计,既作为离线评估的倾向性模型,也用于诊断系统的自适应性能。在四个大规模真实数据集上验证了该框架的有效性,新策略持续优于原始日志策略。结果表明,有效教学策略可在数秒内生成并可视化,为无需额外数据收集的自适应学习系统提供可扩展的改进路径。

原文摘要 · Abstract (English)

Optimizing instructional policies in Intelligent Tutoring Systems (ITS) typically requires costly online experimentation or student simulators that may fail to capture real-world dynamics. This paper introduces an offline contextual bandit framework that learns new adaptive policies directly from logged interaction data. By mapping student-item interactions onto a continuous latent proficiency-difficulty scale using a Rasch model, we cast the tutoring process as a continuous stochastic bandit problem. We propose a novel reward function designed to optimize ''flow'' by balancing task challenge with student success. Our approach includes a round-specific behavior policy estimation that serves as both a propensity model for off-policy evaluation and a diagnostic tool for ITS adaptivity. We demonstrate the efficacy of this framework across four large-scale real-world datasets, achieving consistent policy improvements over the logged behavior policy. The results show that effective instructional policies can be learned and visualized within seconds of computation, providing a scalable path for improving adaptive learning systems without further data collection.

自适应学习离线强化学习智能辅导系统教育人工智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。