让AI导师学会按教育理论思考,提升教学逻辑性
Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
- 用教育理论引导模型内部推理,而非通用指令
- 奖励模型的思考过程,显著提升教学逻辑性
- 适合教育AI、智能辅导系统研发者参考
大语言模型在智能辅导系统中的应用日益广泛,但针对教育场景优化的研究仍有限。现有强化学习方法仅关注可观察的回答,忽视模型内部思考过程。本文提出PedagogicalRL-Thinking框架,通过两种新方法实现教育导向的推理对齐:(1) 教育学理论驱动的推理提示,使用领域特定教育理论指导内部思考;(2) 思考奖励机制,显式评估并强化模型推理轨迹的教育质量。实验表明,基于教育理论的提示优于通用提示,且与思考奖励结合效果最佳。仅在数学辅导对话上训练的模型,在未见教育基准测试中表现提升,同时保持原始模型的事实知识。定量与定性分析显示,该方法使推理轨迹更具系统性,提升教学推理与结构化决策能力。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as intelligent tutoring systems, yet research on optimizing LLMs specifically for educational contexts remains limited. Recent works have proposed reinforcement learning approaches for training LLM tutors, but these methods focus solely on optimizing visible responses while neglecting the model's internal thinking process. We introduce PedagogicalRL-Thinking, a framework that extends pedagogical alignment to reasoning LLMs in education through two novel approaches: (1) Pedagogical Reasoning Prompting, which guides internal reasoning using domain-specific educational theory rather than generic instructions; and (2) Thinking Reward, which explicitly evaluates and reinforces the pedagogical quality of the model's reasoning traces. Our experiments reveal that domain-specific, theory-grounded prompting outperforms generic prompting, and that Thinking Reward is most effective when combined with pedagogical prompting. Furthermore, models trained only on mathematics tutoring dialogues show improved performance on educational benchmarks not seen during training, while preserving the base model's factual knowledge. Our quantitative and qualitative analyses reveal that pedagogical thinking reward produces systematic reasoning trace changes, with increased pedagogical reasoning and more structured instructional decision-making in the tutor's thinking process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。