arXiv:2604.04251cs.AIcs.CY2026-04

用知识掌握度约束教学策略,防止系统只追求学生互动而忽视真实学习。

MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems

论文配图:MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems
图 1 · 摘自论文原文
  • 根据学生掌握程度动态开放教学内容,确保学习路径符合认知逻辑。
  • 在两个平台测试中,知识掌握提升分别提高18.3%和54.0%。
  • 适合需要安全、可信赖自适应教学系统的教育AI研发者。

智能辅导系统越来越多地依赖强化学习实现个性化教学,但仅优化可观测的参与信号会导致学习行为与真实知识获取脱节。分析两个部署平台超过2100万次学生交互数据发现,在Junyi Academy上,26.5%的互动事件无对应掌握度提升(72,758名学生),在XES3G5M上为3.1%(14,453名学生,NeurIPS 2023),表明该现象在真实教育科技系统中广泛存在。本文提出基于掌握度约束的策略优化方法(MC-CPO),其核心思想是:只有当前置知识达到掌握阈值时,新概念才可被呈现,从而自然扩展可选教学动作空间。通过结构化设计实现教学安全约束,具备前提安全性的形式保证、原对偶收敛性及严格优于事后过滤的特性。MC-CPO是唯一在所有条件下均降低奖励滥用严重性的方法。相较于所有基线模型,平均每回合掌握度提升在Junyi Academy上增加18.3%,在XES3G5M上提升54.0%,同时保持相当的参与度表现。结果表明,结构化约束建模是构建可信赖自适应教学策略的坚实基础。

原文摘要 · Abstract (English)

Intelligent tutoring systems increasingly rely on reinforcement learning to personalise instruction, yet optimising for observable engagement signals can systematically decouple learner activity from genuine knowledge acquisition. Analysing over 21 million student interactions across two deployed platforms, we find engagement events without corresponding mastery gains occur in 26.5% of interactions on Junyi Academy (72,758 students) and 3.1% on XES3G5M (14,453 students, NeurIPS 2023), confirming this pattern is directly observable in deployed educational technology at scale. We introduce Mastery-Conditioned Constrained Policy Optimisation (MC-CPO), a reinforcement learning framework that addresses this problem structurally. MC-CPO conditions the admissible instructional action space on learner mastery state: a concept becomes available only when prerequisite knowledge meets a mastery threshold, yielding an action space that expands naturally as learners acquire knowledge. Pedagogical safety constraints are enforced by construction, with formal guarantees of structural prerequisite safety, primal-dual convergence, and strict dominance over post-hoc filtering. MC-CPO is the only method to reduce reward hacking severity across all conditions. Mean per-episode mastery gain increases by 18.3% on Junyi Academy and 54.0% on XES3G5M relative to all baselines, while competitive engagement performance is maintained. These results support structural constraint modelling as a principled foundation for safer adaptive instructional policies in deployed tutoring systems.

智能辅导强化学习教育AI安全策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。