arXiv:2507.19748cs.CL2025-07被引 2

通过多阶段训练提升大模型数学推理能力,开源效果领先。

JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models

  • 分三阶段优化:预训练、指令微调、思维链强化学习
  • 在32K上下文下解决竞赛级数学题,性能超O1-mini和GPT-4o
  • 数据质量高,使用模型验证筛选2100亿词元高质量语料

数学推理是通用人工智能的核心,也是评估大语言模型能力的关键基准。尽管先进模型展现出潜力,但在需要深层概念理解与复杂多步推演的难题上仍表现不佳。为此,我们提出JT-Math-8B,一个基于系统化多阶段优化框架的开源模型系列,包含基础版、指令版和思维版。预训练语料为2100亿词元,通过模型驱动的数据管道确保质量与多样性。指令版经监督微调(SFT)与基于GRPO的强化学习优化,以生成简洁准确答案。思维版采用长链式思维(Long CoT)方法,结合SFT与新型多阶段强化学习课程,逐步提升任务难度与上下文长度至32K tokens。JT-Math-8B在同类开源模型中达到最先进水平,超越OpenAI O1-mini和GPT-4o,在竞赛级数学任务中表现优异。

原文摘要 · Abstract (English)

Mathematical reasoning is a cornerstone of artificial general intelligence and a primary benchmark for evaluating the capabilities of Large Language Models (LLMs). While state-of-the-art models show promise, they often falter when faced with complex problems that demand deep conceptual understanding and intricate, multi-step deliberation. To address this challenge, we introduce JT-Math-8B, a series of open-source models comprising base, instruct, and thinking versions, built upon a systematic, multi-stage optimization framework. Our pre-training corpus is a high-quality, 210B-token dataset curated through a dedicated data pipeline that uses model-based validation to ensure quality and diversity. The Instruct Model is optimized for direct, concise answers through Supervised Fine-Tuning (SFT) and a GRPO-based reinforcement learning (RL) method. The Thinking Model is trained for complex problem-solving using a Long Chain-of-Thought (Long CoT) approach, combining SFT with a novel, multi-stage RL curriculum that progressively increases task difficulty and context length up to 32K tokens. JT-Math-8B achieves state-of-the-art results among open-source models of similar size, surpassing prominent models like OpenAI's O1-mini and GPT-4o , and demonstrating superior performance on competition-level mathematics.

数学推理大模型强化学习开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。