arXiv:2504.00829cs.CL2025-04被引 17

按难度分阶段强化学习,让小模型推理能力大幅提升

How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study

  • 按题目难易分级选数据,优化训练效率
  • 1.5B模型在AIME-2024达42.3%准确率
  • 适合想提升模型逻辑推理能力的研究者

提升大语言模型(LLMs)的推理能力在人工智能研究中仍面临效率与可扩展性的根本挑战。本文通过系统实验,验证了难度感知的分阶段强化学习策略能显著提升模型推理性能。通过根据明确的难度等级选择训练数据,有效增强了强化学习优化效果;同时提出逐步暴露模型于更复杂任务的分阶段训练方法,进一步增强推理能力。跨领域联合训练数学推理与代码生成任务时展现出显著优势。结果显示,所提方法使1.5亿参数模型在AIME-2024基准上达到42.3%准确率,在MATH-500上达89.5%。这些结果证明该方法能有效提升LLMs的推理水平。我们将开源相关数据集至GitHub和Hugging Face。

原文摘要 · Abstract (English)

Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research. This paper presents a rigorous experimental investigation into how difficulty-aware staged reinforcement learning (RL) strategies can substantially improve LLM reasoning performance. Through systematic analysis, we demonstrate that strategically selecting training data according to well-defined difficulty levels markedly enhances RL optimization. Moreover, we introduce a staged training methodology, progressively exposing models to increasingly challenging tasks, further amplifying reasoning capabilities. Our findings reveal significant cross-domain benefits when simultaneously training models on mathematical reasoning and code generation tasks. Notably, our proposed approach enables a 1.5B parameter model to achieve an accuracy of 42.3\% on the AIME-2024 benchmark, 89.5\% on the MATH-500 benchmark. These results underscore the efficacy of our method in advancing the reasoning proficiency of LLMs. We will open-source our datasets on GitHub and Hugging Face.

强化学习推理能力分阶段训练数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。