arXiv:2604.00258cs.LGcs.AI2026-04

用分层框架分析学生错误行为,自动识别学习进展与策略改进。

Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards

论文配图:Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
图 1 · 摘自论文原文
  • 构建分层框架,从不完美学习轨迹中提取多层级意图信号。
  • 能区分临时错误与有效策略演进,准确预测教学决策。
  • 适合教育智能系统研发者,提升个性化学习支持能力。

尽管模仿学习在在线教育环境中已展现潜力,可直接从学生交互中生成有效教学策略,但现有方法通常依赖固定奖励下的最优或近似最优示范。现实中,学生行为常表现为不完美且动态变化:探索、出错、修正策略并逐步明确目标。本文认为,不完美的学生示范并非噪声,而是蕴含结构信息的信号,前提是其质量可被排序。为此提出HALIDE(分层模仿学习从动态演化奖励的不完美示范),不仅利用次优示范,还在分层框架中对它们进行排序。该模型在多个抽象层次上建模学生行为,从次优动作推断高层意图与策略,并显式捕捉学生奖励函数随时间的演变。通过将示范质量整合到分层奖励推断中,HALIDE可区分暂时性错误与向更高学习目标的有效推进。实验表明,相比依赖最优轨迹、固定奖励或未排序不完美示范的方法,HALIDE更准确地预测学生教学决策。

原文摘要 · Abstract (English)

While apprenticeship learning has shown promise for inducing effective pedagogical policies directly from student interactions in e-learning environments, most existing approaches rely on optimal or near-optimal expert demonstrations under a fixed reward. Real-world student interactions, however, are often inherently imperfect and evolving: students explore, make errors, revise strategies, and refine their goals as understanding develops. In this work, we argue that imperfect student demonstrations are not noise to be discarded, but structured signals-provided their relative quality is ranked. We introduce HALIDE, Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards, which not only leverages sub-optimal student demonstrations, but ranks them within a hierarchical learning framework. HALIDE models student behavior at multiple levels of abstraction, enabling inference of higher-level intent and strategy from suboptimal actions while explicitly capturing the temporal evolution of student reward functions. By integrating demonstration quality into hierarchical reward inference,HALIDE distinguishes transient errors from suboptimal strategies and meaningful progress toward higher-level learning goals. Our results show that HALIDE more accurately predicts student pedagogical decisions than approaches that rely on optimal trajectories, fixed rewards, or unranked imperfect demonstrations.

模仿学习教育智能分层建模动态奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。