arXiv:2601.14686cs.AIcs.LG2026-01被引 2

用多目标强化学习让大模型推荐的学习路径更符合教学目标。

IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

  • 基于指标的分组相对策略优化,自动平衡多个教学目标。
  • 在两个数据集上优于主流强化学习与大模型基线方法。
  • 适合需要个性化学习路径设计的教育科技开发者。

学习路径推荐(LPR)旨在生成个性化学习序列,在尊重教学原则和操作约束的前提下最大化长期学习效果。尽管大语言模型(LLM)具备丰富的语义理解能力,但在长时序推荐中仍面临三大挑战:(i)在稀疏延迟反馈下与教学目标(如最近发展区ZPD)存在偏差;(ii)专家示范数据稀缺且成本高;(iii)学习效果、难度调度、长度可控性与轨迹多样性之间的多目标冲突。为此,我们提出IB-GRPO(基于指标的分组相对策略优化),一种指导式对齐方法。通过遗传算法搜索与教师强化学习代理构建混合专家示范,并以监督微调方式预热LLM。在此基础上,设计会话内ZPD对齐得分用于难度调度。IB-GRPO采用$ I_{ε+} $占优指标计算多目标的组相对优势,避免手动加权,提升帕累托权衡表现。在ASSIST09与Junyi数据集上,基于Qwen2.5-7B模型与KES模拟器的实验显示,该方法持续优于代表性强化学习与大模型基线。

原文摘要 · Abstract (English)

Learning Path Recommendation (LPR) aims to generate personalized sequences of learning items that maximize long-term learning effect while respecting pedagogical principles and operational constraints. Although large language models (LLMs) offer rich semantic understanding for free-form recommendation, applying them to long-horizon LPR is challenging due to (i) misalignment with pedagogical objectives such as the Zone of Proximal Development (ZPD) under sparse, delayed feedback, (ii) scarce and costly expert demonstrations, and (iii) multi-objective interactions among learning effect, difficulty scheduling, length controllability, and trajectory diversity. To address these issues, we propose IB-GRPO (Indicator-Based Group Relative Policy Optimization), an indicator-guided alignment approach for LLM-based LPR. To mitigate data scarcity, we construct hybrid expert demonstrations via Genetic Algorithm search and teacher RL agents and warm-start the LLM with supervised fine-tuning. Building on this warm-start, we design a within-session ZPD alignment score for difficulty scheduling. IB-GRPO then uses the $I_{ε+}$ dominance indicator to compute group-relative advantages over multiple objectives, avoiding manual scalarization and improving Pareto trade-offs. Experiments on ASSIST09 and Junyi using the KES simulator with a Qwen2.5-7B backbone show consistent improvements over representative RL and LLM baselines.

学习路径推荐多目标优化大模型应用教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。