用进化强化学习打造能动态互动的智能导师,提升跨学科思维能力。
Evolutionary Reinforcement Learning based AI tutor for Socratic Interdisciplinary Instruction
- 基于知识图谱建模学生认知状态,实现动态交互
- 分层奖励机制缓解长期目标奖励稀疏问题
- 结合演化算法与PPO,兼顾策略多样性和精准优化
现代STEM教育需培养知识整合、批判性思维和创造力等高阶认知能力,要求从被动传授转向主动苏格拉底式建构。尽管大语言模型(LLMs)在跨学科教育中具有潜力,但现有基于提示工程(PE)、监督微调(SFT)或标准强化学习(RL)的方法难以支持此范式。主要受限于三方面:无法动态建模隐含的学生认知状态;长期教育目标导致奖励严重稀疏且延迟;依赖行为克隆易引发策略坍塌,缺乏策略多样性。为此,我们首次将苏格拉底跨学科教学问题(SIIP)形式化为结构化部分可观测马尔可夫决策过程(POMDP),要求同时实现全局探索与精细策略优化。提出ERL4SIIP框架:(1) 基于STEM知识图谱的动态学生模拟器,用于建模潜在认知状态;(2) 分层奖励机制,将长期目标分解为密集信号;(3) 融合LoRA-分割优化策略,以演化算法进行种群级全局搜索,配合PPO进行局部梯度上升。
原文摘要 · Abstract (English)
Cultivating higher-order cognitive abilities -- such as knowledge integration, critical thinking, and creativity -- in modern STEM education necessitates a pedagogical shift from passive knowledge transmission to active Socratic construction. Although Large Language Models (LLMs) hold promise for STEM Interdisciplinary education, current methodologies employing Prompt Engineering (PE), Supervised Fine-tuning (SFT), or standard Reinforcement Learning (RL) often fall short of supporting this paradigm. Existing methods are hindered by three fundamental challenges: the inability to dynamically model latent student cognitive states; severe reward sparsity and delay inherent in long-term educational goals; and a tendency toward policy collapse lacking strategic diversity due to reliance on behavioral cloning. Recognizing the unobservability and dynamic complexity of these interactions, we formalize the Socratic Interdisciplinary Instructional Problem (SIIP) as a structured Partially Observable Markov Decision Process (POMDP), demanding simultaneous global exploration and fine-grained policy refinement. To this end, we propose ERL4SIIP, a novel Evolutionary Reinforcement Learning (ERL) framework specifically tailored for this domain. ERL4SIIP integrates: (1) a dynamic student simulator grounded in a STEM knowledge graph for latent state modeling; (2) a Hierarchical Reward Mechanism that decomposes long-horizon goals into dense signals; and (3) a LoRA-Division based optimization strategy coupling evolutionary algorithms for population-level global search with PPO for local gradient ascent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。