arXiv:2605.29582cs.LGcs.CL2026-05

用强化学习训练能循序渐进引导学生的智能导师。

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

论文配图:PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
图 1 · 摘自论文原文
  • 构建可控制的学生模拟器,分离认知状态与答题行为。
  • 设计兼顾教学质量和答案正确的多目标奖励模型。
  • 适合教育AI研究者和智能辅导系统开发者使用。

大型语言模型在教育辅导方面展现出巨大潜力。现有方法通常训练模型直接解题并给出正确答案,但这种以解题为中心的模式忽略了有效辅导的关键需求:逐步引导以及多轮互动中协调多种教学目标。由于学生知识水平差异大、教学效果不仅取决于答案正确性,且多目标协调在交互中极具挑战,因此开发此类导师仍困难重重。为此,我们提出PEARL框架——一种用于训练苏格拉底式辅导代理的教育对齐强化学习方法。首先,引入可控学生模拟器,将潜在认知状态与响应生成解耦,实现不同能力与误解的模拟;其次,构建教育对齐的奖励模型,联合评估教学质量和答案正确性;最后,提出稳定的多目标强化学习方法,在训练过程中平衡冲突的教学目标。跨多个基准测试的实验表明,PEARL性能优于开源辅导系统,并媲美领先专有LLM。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show strong potential as educational tutors. Existing approaches typically train them to solve problems and provide correct answers, but this problem-solving-centered paradigm overlooks key requirements of effective tutoring: progressive guidance and the coordination of multiple pedagogical objectives across multi-turn interactions. Developing such tutors remains challenging because student behavior varies substantially with individual knowledge states, pedagogical effectiveness depends on multiple factors beyond final-answer correctness, and coordinating these objectives over tutor-student interactions is inherently difficult. To address these challenges, we propose PEARL, a PEdagogically Aligned Reinforcement Learning framework for training Socratic tutoring agents. First, we introduce a controllable student simulator that disentangles latent cognitive states from response generation, enabling simulation of diverse abilities and misconceptions. Second, we develop a pedagogically aligned reward model that jointly assesses pedagogical quality and objective correctness. Finally, we propose a stable multi-objective reinforcement learning approach that balances competing pedagogical objectives during tutor training. Experiments across multiple benchmarks show that PEARL performs competitively against tutoring-specific open-source systems and leading proprietary LLMs.

智能辅导强化学习教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。