arXiv:2508.08909cs.AI2025-08

用更少资源训练出强推理模型,数学能力达40%准确率

Compass-Thinker-7B Technical Report

  • 基于开源模型设计强化学习流水线,分阶段释放推理潜力
  • 在3万道数学题上训练,AIME2024测试达40%准确率
  • 为大模型强化学习提供低成本实验方案,适合算法研究者

近期类似R1-Zero的研究进一步证明,推理扩展使大语言模型获得前所未有的推理能力,强化学习是激发其复杂推理的核心技术。然而,在超大规模模型上直接进行强化学习实验成本高昂且风险大。本文提出Compass-Thinker-7B模型,旨在以较低计算资源和成本探索强化学习潜力,并为更大模型的强化学习方法提供参考。该模型基于开源模型,通过专门设计的强化学习流水线训练而成。我们构建了一个包含3万道可验证数学问题的数据集用于训练。通过在不同阶段配置具有不同难度分布的数据与训练设置,逐步释放模型潜能并提升训练效率。大量评估表明,Compass-Thinker-7B具备卓越的推理潜力,在数学任务上表现优于同规模的强化学习模型。尤其在具有挑战性的AIME2024评测中,达到40%准确率。

原文摘要 · Abstract (English)

Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning is the core technology to elicit its complex reasoning. However, conducting RL experiments directly on hyperscale models involves high computational costs and resource demands, posing significant risks. We propose the Compass-Thinker-7B model, which aims to explore the potential of Reinforcement Learning with less computational resources and costs, and provides insights for further research into RL recipes for larger models. Compass-Thinker-7B is trained from an open source model through a specially designed Reinforcement Learning Pipeline. We curate a dataset of 30k verifiable mathematics problems for the Reinforcement Learning Pipeline. By configuring data and training settings with different difficulty distributions for different stages, the potential of the model is gradually released and the training efficiency is improved. Extensive evaluations show that Compass-Thinker-7B possesses exceptional reasoning potential, and achieves superior performance on mathematics compared to the same-sized RL model. Especially in the challenging AIME2024 evaluation, Compass-Thinker-7B achieves 40% accuracy.

强化学习数学推理小模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。