用程序生成1230万组数学题解,提升大模型推理能力
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
- 通过程序生成数学题与解法,保证内容可执行
- 生成1230万组题解对,经双重验证确保正确性
- 适合需要高质量数学训练数据的研究者
提升大语言模型的数学推理能力需要高质量训练数据,但传统方法在可扩展性、成本和数据可靠性方面面临挑战。为此,我们提出一种基于程序的合成框架,系统生成具有保障多样性和复杂性的高质量数学语料库。该框架结合数学知识体系与领域专用工具,生成可执行程序,并将其转换为自然语言的问题-解答对,再通过双侧验证机制检验答案正确性与问题-程序一致性。共生成1230万组此类问题求解三元组。实验表明,使用该数据微调的模型在多个基准测试中显著提升推理能力,达到当前最优水平,验证了该合成方法的有效性。
原文摘要 · Abstract (English)
Enhancing the mathematical reasoning of large language models (LLMs) demands high-quality training data, yet conventional methods face critical challenges in scalability, cost, and data reliability. To address these limitations, we propose a novel program-assisted synthesis framework that systematically generates a high-quality mathematical corpus with guaranteed diversity, complexity, and correctness. This framework integrates mathematical knowledge systems and domain-specific tools to create executable programs. These programs are then translated into natural language problem-solution pairs and vetted by a bilateral validation mechanism that verifies solution correctness against program outputs and ensures program-problem consistency. We have generated 12.3 million such problem-solving triples. Experiments demonstrate that models fine-tuned on our data significantly improve their inference capabilities, achieving state-of-the-art performance on several benchmark datasets and showcasing the effectiveness of our synthesis approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。