arXiv:2604.06385cs.CL2026-04

用强化学习和监督微调让开源大模型变教学专家,效果超更大商用模型。

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning

  • 分三阶段优化:先强化学习攻坚难题,再用其生成高质量训练数据,最后可再强化一次。
  • 320亿参数模型在教学知识基准上达新纪录,超越更大商用模型如Gemini-3 Pro。
  • 适合教育AI研发者,兼具透明性、可定制性与低成本优势。

我们提出一种结合强化学习(RL)与监督微调(SFT)的多阶段优化策略,以提升大型语言模型(LLMs)的教学能力,具体表现为EduQwen 32B-RL1、EduQwen 32B-SFT及可选的第三阶段模型EduQwen 32B-SFT-RL2:(1)RL优化采用渐进难度训练,聚焦挑战性例题,并使用扩展推理回溯;(2)后续SFT阶段利用已训练的RL模型合成高质量训练数据,采用难度加权采样;(3)可选的第二轮RL优化。这些基于密集Qwen3-32B骨干的开源教学专用模型,在跨领域教学知识(CDPK)基准上达到显著高准确率,刷新交互式教学基准排行榜(Interactive Pedagogy Benchmark Leaderboard)新纪录,显著超越更大规模的专有系统如Gemini-3 Pro。320亿参数的密集模型证明,领域专项优化可使中等规模开源模型成为真正的教学专家,优于更大通用系统,同时保持教育AI部署所需的透明性、可定制性和成本效益。

原文摘要 · Abstract (English)

We present an innovative multi-stage optimization strategy combining reinforcement learning (RL) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of large language models (LLMs), as illustrated by EduQwen 32B-RL1, EduQwen 32B-SFT, and an optional third-stage model EduQwen 32B-SFT-RL2: (1) RL optimization that implements progressive difficulty training, focuses on challenging examples, and employs extended reasoning rollouts; (2) a subsequent SFT phase that leverages the RL-trained model to synthesize high-quality training data with difficulty-weighted sampling; and (3) an optional second round of RL optimization. EduQwen 32B-RL1, EduQwen 32B-SFT, and EduQwen 32B-SFT-RL2 are an application-driven family of open-source pedagogical LLMs built on a dense Qwen3-32B backbone. These models remarkably achieve high enough accuracy on the Cross-Domain Pedagogical Knowledge (CDPK) Benchmark to establish new state-of-the-art (SOTA) results across the interactive Pedagogy Benchmark Leaderboard and surpass significantly larger proprietary systems such as the previous benchmark leader Gemini-3 Pro. These dense 32-billion-parameter models demonstrate that domain-specialized optimization can transform mid-sized open-source LLMs into true pedagogical domain experts that outperform much larger general-purpose systems, while preserving the transparency, customizability, and cost-efficiency required for responsible educational AI deployment.

教学大模型强化学习开源模型微调策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。