arXiv:2412.15084cs.CLcs.AI2024-12被引 79

AceMath通过强化训练与奖励模型,显著提升数学推理能力。

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

  • 采用分阶段微调策略,先通用后专精,优化数学解题能力。
  • 在多个数学基准上,平均得分超越GPT-4o和Claude-3.5 Sonnet。
  • 开源模型与评测基准,适合数学推理研究者使用。

本文提出AceMath,一套面向复杂数学问题求解的前沿模型,以及高效可靠的奖励模型,可准确评估生成解法并识别正确答案。为构建指令微调数学模型,我们设计了监督微调流程:先在通用领域达到良好表现,再针对数学领域使用精心筛选的提示和合成响应进行定向微调。由此产生的AceMath-72B-Instruct在多项指标上显著优于Qwen2.5-Math-72B-Instruct、GPT-4o和Claude-3.5 Sonnet。为构建专用数学奖励模型,我们构建了AceMath-RewardBench,一个覆盖多种题型与难度的综合性评测基准。随后提出系统化方法训练奖励模型。最终的AceMath-72B-RM持续优于现有先进奖励模型。将两者结合后,在数学推理基准上实现最高rm@8平均得分。模型权重、训练数据及评测基准已公开于https://research.nvidia.com/labs/adlr/acemath。

原文摘要 · Abstract (English)

In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifying the correct ones. To develop the instruction-tuned math models, we propose a supervised fine-tuning (SFT) process that first achieves competitive performance across general domains, followed by targeted fine-tuning for the math domain using a carefully curated set of prompts and synthetically generated responses. The resulting model, AceMath-72B-Instruct greatly outperforms Qwen2.5-Math-72B-Instruct, GPT-4o and Claude-3.5 Sonnet. To develop math-specialized reward model, we first construct AceMath-RewardBench, a comprehensive and robust benchmark for evaluating math reward models across diverse problems and difficulty levels. After that, we present a systematic approach to build our math reward models. The resulting model, AceMath-72B-RM, consistently outperforms state-of-the-art reward models. Furthermore, when combining AceMath-72B-Instruct with AceMath-72B-RM, we achieve the highest average rm@8 score across the math reasoning benchmarks. We release model weights, training data, and evaluation benchmarks at: https://research.nvidia.com/labs/adlr/acemath

数学推理奖励模型大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。