arXiv:2505.15607cs.CLcs.AI2025-05EMNLP被引 20

用强化学习让大模型学会教学,不直接给答案而引导解题。

From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning

  • 通过模拟师生互动,用强化学习优化模型教学策略。
  • 70亿参数模型表现媲美更大商用模型,且保留更强推理能力。
  • 可调节教学支持与解题准确率平衡,适合教育类应用开发。

大型语言模型(LLMs)有望改变教育,但其针对直接问答的优化常违背有效教学原则——即应策略性地延迟给出答案。为此,我们提出一种基于在线强化学习(RL)的对齐框架,通过模拟学生-教师交互,快速将LLMs调整为高效导师,强调教学质量和引导式解题,而非直接提供答案。我们利用该方法训练了一个70亿参数的导师模型,无需人工标注,性能接近更大的专有模型LearnLM。引入可控奖励权重,可平衡教学支持与学生解题准确率,实现两者间的帕累托前沿追踪。相比单轮微调基线,本模型更有效地保持了推理能力,并可通过思维标签增强可解释性,揭示模型的教学规划过程。

原文摘要 · Abstract (English)

Large language models (LLMs) can transform education, but their optimization for direct question-answering often undermines effective pedagogy which requires strategically withholding answers. To mitigate this, we propose an online reinforcement learning (RL)-based alignment framework that can quickly adapt LLMs into effective tutors using simulated student-tutor interactions by emphasizing pedagogical quality and guided problem-solving over simply giving away answers. We use our method to train a 7B parameter tutor model without human annotations which reaches similar performance to larger proprietary models like LearnLM. We introduce a controllable reward weighting to balance pedagogical support and student solving accuracy, allowing us to trace the Pareto frontier between these two objectives. Our models better preserve reasoning capabilities than single-turn SFT baselines and can optionally enhance interpretability through thinking tags that expose the model's instructional planning.

教学模型强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。