小模型通过预判难度增强数学推理能力,效果媲美大模型。
Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- 在开源框架上设计难度感知干预策略,提升小模型数学推理。
- 1.5B参数模型在AIME24上达50.0%,超越O1-Preview。
- 适合关注小模型高效训练与可复现性研究的学者。
强化学习的扩展提升了大语言模型的推理能力,但当前顶尖推理模型(如OpenAI O系列、Claude 3系列、DeepMind Gemini 2.5系列、Grok 3系列)的关键技术细节未公开,导致研究社区难以复现其强化学习训练结果。为此,我们基于开源GRPO框架提出早期预览强化学习(EPRLI),引入数学题难度感知干预机制。该方法应用于1.5亿参数的小型语言模型,在标准实验室环境下实现AIME24 50.0%、Math500 89.2%、AMC 77.1%、Minerva 35.3%、OBench 51.9%的准确率,性能超越O1-Preview,并接近O1-mini水平。
原文摘要 · Abstract (English)
Reinforcement learning scaling enhances the reasoning capabilities of large language models, with reinforcement learning serving as the key technique to draw out complex reasoning. However, key technical details of state-of-the-art reasoning LLMs, such as those in the OpenAI O series, Claude 3 series, DeepMind's Gemini 2.5 series, and Grok 3 series, remain undisclosed, making it difficult for the research community to replicate their reinforcement learning training results. Therefore, we start our study from an Early Preview Reinforcement Learning (EPRLI) algorithm built on the open-source GRPO framework, incorporating difficulty-aware intervention for math problems. Applied to a 1.5B-parameter LLM, our method achieves 50.0% on AIME24, 89.2% on Math500, 77.1% on AMC, 35.3% on Minerva, and 51.9% on OBench, superpass O1-Preview and is comparable to O1-mini within standard school-lab settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。