通过自我迭代提升数学推理能力,打造中英双语数学专家模型。
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

- 构建自进化体系:从预训练到推理全程使用奖励模型驱动数据与模型迭代优化。
- 在10个中英文数学数据集上表现优异,最高达AIME24竞赛级解题水平。
- 适合需要高精度数学推理的科研、教育及竞赛场景使用。
本文介绍一系列专精于数学的大型语言模型:Qwen2.5-Math 及 Qwen2.5-Math-Instruct-1.5B/7B/72B。其核心创新在于将自改进理念贯穿于预训练、后训练到推理的全流程:(1)预训练阶段,利用 Qwen2-Math-Instruct 生成大规模高质量数学数据;(2)后训练阶段,通过大量采样构建奖励模型(RM),并用于监督微调(SFT)中的数据迭代演化。更强的 SFT 模型可反向优化 RM,形成闭环迭代。最终在 SFT 模型上应用终极 RM 进行强化学习,得到 Qwen2.5-Math-Instruct。此外,在推理阶段,亦采用该 RM 引导采样以优化输出。模型支持中英文,具备链式思维(CoT)与工具融合推理(TIR)能力。我们在涵盖小学至数学竞赛级别的10个数学数据集(如 GSM8K、MATH、GaoKao、AMC23、AIME24)上进行评估。
原文摘要 · Abstract (English)
In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolution of data in supervised fine-tuning (SFT). With a stronger SFT model, it's possible to iteratively train and update the RM, which in turn guides the next round of SFT data iteration. On the final SFT model, we employ the ultimate RM for reinforcement learning, resulting in the Qwen2.5-Math-Instruct. (3) Furthermore, during the inference stage, the RM is used to guide sampling, optimizing the model's performance. Qwen2.5-Math-Instruct supports both Chinese and English, and possess advanced mathematical reasoning capabilities, including Chain-of-Thought (CoT) and Tool-Integrated Reasoning (TIR). We evaluate our models on 10 mathematics datasets in both English and Chinese, such as GSM8K, MATH, GaoKao, AMC23, and AIME24, covering a range of difficulties from grade school level to math competition problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。