arXiv:2602.20527cs.LGcs.AI2026-02被引 3

用少量学生学习轨迹,自动学习并预测教学策略。

A Generalized Apprenticeship Learning Framework for Capturing Evolving Student Pedagogical Strategies

  • 基于专家示范动态推断多维奖励函数,捕捉学习过程演化。
  • 仅用18条历史轨迹,即达AUC 0.899、Jaccard 0.653的预测精度。
  • 适合教育智能系统中的个性化教学策略生成与优化。

近年来,强化学习(RL)与深度强化学习(DRL)在智能辅导系统等教育环境中取得显著进展。然而,其在教育技术中的广泛应用受限于样本效率低及奖励函数设计困难等问题。相比之下,示范学习(AL)可通过少量专家示范推断其隐含奖励函数,并生成可泛化、复现最优行为的决策策略。本文提出一种广义示范学习框架THEMES,通过捕捉学生学习过程中多维度奖励函数的动态演化,实现对有效教学策略的建模。在六种先进基线方法上的评估表明,THEMES表现优异,仅使用上一学期18条学习轨迹,即可在下一学期学生决策预测中达到AUC 0.899和Jaccard 0.653的高精度,展现出作为高效教学策略生成工具的巨大潜力。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) and Deep Reinforcement Learning (DRL) have advanced rapidly in recent years and have been successfully applied to e-learning environments like intelligent tutoring systems (ITSs). Despite great success, the broader application of DRL to educational technologies has been limited due to major challenges such as sample inefficiency and difficulty designing the reward function. In contrast, Apprenticeship Learning (AL) uses a few expert demonstrations to infer the expert's underlying reward functions and derive decision-making policies that generalize and replicate optimal behavior. In this work, we leverage a generalized AL framework, THEMES, to induce effective pedagogical policies by capturing the complexities of the expert student learning process, where multiple reward functions may dynamically evolve over time. We evaluate the effectiveness of THEMES against six state-of-the-art baselines, demonstrating its superior performance and highlighting its potential as a powerful alternative for inducing effective pedagogical policies and show that it can achieve high performance, with an AUC of 0.899 and a Jaccard of 0.653, using only 18 trajectories of a previous semester to predict student pedagogical decisions in a later semester.

教育AI示范学习策略生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。