通过缺陷感知框架提升大模型数学能力,效果优于现有方法。
WarriorMath: Enhancing the Mathematical Ability of Large Language Models with a Defect-aware Framework
- 用多专家协作生成并优化模型易错的数学题
- 在6个基准上平均提升12.57%,达新最优
- 适合想提升模型数学推理能力的研究者
大型语言模型在解决数学问题上表现优异,但其性能常受限于高质量、多样化训练数据的缺乏。现有方法多通过重述或难度递进扩充数据集,却忽视了模型的具体失败模式,导致合成题目大多为模型已能解答的内容,难以带来显著提升。为此,我们提出WarriorMath——一种缺陷感知的数学问题求解框架,融合定向数据合成与渐进式训练。在合成阶段,多个专家级LLM协同生成、评估并优化题目,识别出模型无法解答的问题,并通过专家反馈迭代改进,生成高质量、缺陷感知的训练数据。在训练阶段,设计渐进学习框架,使用针对模型弱点的渐进挑战数据进行迭代微调。在六个数学基准上的实验表明,WarriorMath平均性能优于强基线12.57%,达到新的最佳水平。结果证明,缺陷感知的多专家框架在提升数学能力方面具有显著有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in solving mathematical problems, yet their performance is often limited by the availability of high-quality, diverse training data. Existing methods focus on augmenting datasets through rephrasing or difficulty progression but overlook the specific failure modes of LLMs. This results in synthetic questions that the model can already solve, providing minimal performance gains. To address this, we propose WarriorMath, a defect-aware framework for mathematical problem solving that integrates both targeted data synthesis and progressive training. In the synthesis stage, we employ multiple expert LLMs in a collaborative process to generate, critique, and refine problems. Questions that base LLMs fail to solve are identified and iteratively improved through expert-level feedback, producing high-quality, defect-aware training data. In the training stage, we introduce a progressive learning framework that iteratively fine-tunes the model using increasingly challenging data tailored to its weaknesses. Experiments on six mathematical benchmarks show that WarriorMath outperforms strong baselines by 12.57% on average, setting a new state-of-the-art. Our results demonstrate the effectiveness of a defect-aware, multi-expert framework for improving mathematical ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。