arXiv:2511.15032cs.LGcs.AI2025-11

用强化学习模拟动态课堂,通过干预平衡学生状态估计与教学干扰。

Simulated Human Learning in a Dynamic, Partially-Observed, Time-Series Environment

  • 设计可变探测干预的时序仿真环境,模拟真实教学场景。
  • 探测干预使学生状态估计更准,但需权衡信息收益与干扰成本。
  • 强化学习与启发式策略均灵活适应学生分布变化,但后者对困难班级效果更好。

尽管智能辅导系统(ITS)能利用过往学生数据个性化教学,每位新学生仍具独特性。教育问题本质复杂,因学习过程仅部分可观测。为此,我们构建了一个动态、时序性的仿真教室环境,包含师生互动(如辅导、讲座、考试)。特别地,设计了不同探测强度的干预机制以获取更多信息。我们开发了结合个体状态学习与群体信息利用的强化学习型ITS,通过探测干预降低学生状态估计难度,但需权衡探测频率带来的成本与收益。对比标准强化学习算法与多种启发式规则,发现两者结果相似,但探测干预显著提升性能。随着隐藏信息增加,问题难度上升;允许探测干预带来明显增益。在不同学生群体分布下,两类策略均表现灵活,但强化学习对较难班级支持不足。测试不同课程结构发现,非探测策略在测验和期中考核结构下表现优于仅期末考核结构,凸显持续信息反馈的价值。

原文摘要 · Abstract (English)

While intelligent tutoring systems (ITSs) can use information from past students to personalize instruction, each new student is unique. Moreover, the education problem is inherently difficult because the learning process is only partially observable. We therefore develop a dynamic, time-series environment to simulate a classroom setting, with student-teacher interventions - including tutoring sessions, lectures, and exams. In particular, we design the simulated environment to allow for varying levels of probing interventions that can gather more information. Then, we develop reinforcement learning ITSs that combine learning the individual state of students while pulling from population information through the use of probing interventions. These interventions can reduce the difficulty of student estimation, but also introduce a cost-benefit decision to find a balance between probing enough to get accurate estimates and probing so often that it becomes disruptive to the student. We compare the efficacy of standard RL algorithms with several greedy rules-based heuristic approaches to find that they provide different solutions, but with similar results. We also highlight the difficulty of the problem with increasing levels of hidden information, and the boost that we get if we allow for probing interventions. We show the flexibility of both heuristic and RL policies with regards to changing student population distributions, finding that both are flexible, but RL policies struggle to help harder classes. Finally, we test different course structures with non-probing policies and we find that our policies are able to boost the performance of quiz and midterm structures more than we can in a finals-only structure, highlighting the benefit of having additional information.

强化学习智能辅导动态环境状态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。