用强化学习训练的理科大模型,解题更准且省 token。
Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning

- 通过强化学习+逐步扩大的推理群体,提升解题能力
- 在JEE/NEET等考试中表现优于基线模型,最多少用64%输出词数
- 适合需要高精度、结构化解题的竞赛备考场景
JEE和NEET等顶尖理工科考试要求多步符号推理、精确数值计算及深层概念理解。现有大语言模型虽在通用推理任务上表现良好,但难以规模化部署,无法满足百万学生对领域专用、结构一致解题的需求。我们提出Aryabhata 2,一个面向竞赛级STEM推理的专注型语言模型,基于GPT-OSS-20B通过强化学习后训练实现。利用PhysicsWallah内部题库构建高质量训练课程,采用可验证奖励机制进行强化学习训练。训练结合长期强化学习与逐步扩大的推理群体规模以增强探索能力。在JEE Main、JEE Advanced、NEET,以及AIME、HMMT、MMLU-Pro、MMLU-Redux 2.0、GPQA等分布外推理数据集上评估显示,Aryabhata 2在竞赛推理任务上超越基线模型,同时显著减少输出词数(最高达64%)。
原文摘要 · Abstract (English)
Competitive STEM examinations such as JEE and NEET require multi-step symbolic reasoning, precise numerical computation, and deep conceptual understanding across physics, chemistry, and mathematics. Recent large language models perform strongly on common reasoning benchmarks, yet they remain difficult to deploy at scale, where millions of student doubts demand domain-specific, consistently structured problem solving. We introduce Aryabhata 2, a reasoning-focused language model for competitive STEM examinations, trained via reinforcement-learning post-training. Using PhysicsWallah's internal question banks, we construct a high-quality training curriculum and post-train GPT-OSS-20B through reinforcement learning with verifiable rewards. Training combines prolonged reinforcement learning with broadened exploration via progressively larger rollout group sizes. We evaluate Aryabhata 2 on competitive examination benchmarks, including JEE Main, JEE Advanced, and NEET, as well as out-of-distribution reasoning datasets such as AIME, HMMT, MMLU-Pro, MMLU-Redux 2.0, and GPQA. Results show that Aryabhata 2 outperforms its base model GPT-OSS-20B on competitive STEM reasoning while requiring substantially fewer output tokens (up to 64\% fewer).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。