arXiv:2508.16514cs.LGcs.AI2025-08EMNLP被引 2

通过系统分析合成数据流程,提升大模型数学推理能力。

FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline

  • 构建框架系统评估10种数据合成策略,揭示关键影响因素。
  • 高复杂度问题合成可显著提升多数数学指标表现。
  • 适合关注数学推理、数据合成与模型泛化能力的研究者。

近期通过合成数据提升大模型数学推理能力的工作采用不同设置,难以比较数据合成策略。为此,本文提出FLAMES框架,对10种现有数据合成策略及多个影响因素进行系统研究。实验发现:提升问题复杂度的数据代理在多数数学指标上表现最优;固定生成预算下,保持更高问题覆盖率比仅保留可靠解的问题更重要;基于GSM8K和MATH的合成数据可在竞赛级基准上实现性能提升,展示从易到难的泛化能力。基于这些洞见,我们设计两种新策略以增强域外泛化与鲁棒性。进一步构建了FLAMES数据集,融合新旧策略,在OlympiadBench(+15.7)、CollegeMath(+4.5)、GSMPlus(+6.5)和MATH(+3.1)上优于公开数据集。用FLAMES微调Qwen2.5-Math-7B模型,在MATH上达到81.4%,超越更大模型Llama3 405B、GPT-4o和Claude 3.5 Sonnet。

原文摘要 · Abstract (English)

Recent works improving LLM math reasoning with synthetic data have used unique setups, making comparison of data synthesis strategies impractical. This leaves many unanswered questions about the roles of different factors in the synthetic data pipeline, such as the impact of filtering low-quality problems. To address this gap, we introduce FLAMES, a Framework for LLM Assessment of Math rEasoning Data Synthesis, and perform a systematic study of 10 existing data synthesis strategies and multiple other factors impacting the performance of synthetic math reasoning data. Our FLAMES experiments provide several valuable insights about the optimal balance of difficulty and diversity of synthetic data. First, data agents designed to increase problem complexity lead to best improvements on most math metrics. Second, with a fixed data generation budget, keeping higher problem coverage is more important than keeping only problems with reliable solutions. Third, GSM8K- and MATH-based synthetic data can lead to improvements on competition-level benchmarks, showcasing easy-to-hard generalization. Leveraging insights from our FLAMES experiments, we design two novel data synthesis strategies for improving out-of-domain generalization and robustness. Further, we develop the FLAMES dataset, an effective blend of our novel and existing data synthesis strategies, outperforming public datasets on OlympiadBench (+15.7), CollegeMath (+4.5), GSMPlus (+6.5), and MATH (+3.1). Fine-tuning Qwen2.5-Math-7B on the FLAMES dataset achieves 81.4% on MATH, surpassing larger Llama3 405B, GPT-4o and Claude 3.5 Sonnet.

数学推理数据合成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。