arXiv:2506.06499cs.LGcs.AI2025-06被引 9

用单一模型生成高质量数学题,提升推理模型性能

SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms

  • 基于解题率设计质量-多样性算法,自动生成数学问题
  • 生成超2000万组题目,使模型性能提升最高24%
  • 强调题目难度与多样性对模型泛化能力的差异化作用

大语言模型驱动的合成数据生成已成为提升模型推理能力的有效方法。然而,现有方法多依赖于将大型先进模型压缩到小型学生模型,或使用自然真实问题保证题目质量,限制了在更复杂多样问题领域的可扩展性。为此,我们提出SPARQ:通过质量-多样性算法生成高质且多样化的合成数学问题与解答对,仅需单个模型,以解题率为问题难度的代理指标。从7.5K初始样本出发,生成超过2000万条新问题-解答对。实验表明,按难度筛选生成数据并微调同一模型,可使模型相对性能提升最高达24%。我们还通过消融实验研究了合成数据的数量、质量和多样性对模型泛化的影响:更高的质量(以难度衡量)有助于更好的分布内表现;虽然多样性对分布内性能提升不显著,但筛选更具多样性数据能促进更鲁棒的分布外泛化。此外,我们验证了合成问题存在模型与数据缩放规律,正向促进下游模型泛化。

原文摘要 · Abstract (English)

Large language model (LLM) driven synthetic data generation has emerged as a powerful method for improving model reasoning capabilities. However, most methods either distill large state-of-the-art models into small students or use natural ground-truth problem statements to guarantee problem statement quality. This limits the scalability of these approaches to more complex and diverse problem domains. To address this, we present SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms, a novel approach for generating high-quality and diverse synthetic math problem and solution pairs using only a single model by measuring a problem's solve-rate: a proxy for problem difficulty. Starting from a seed dataset of 7.5K samples, we generate over 20 million new problem-solution pairs. We show that filtering the generated data by difficulty and then fine-tuning the same model on the resulting data improves relative model performance by up to 24\%. Additionally, we conduct ablations studying the impact of synthetic data quantity, quality and diversity on model generalization. We find that higher quality, as measured by problem difficulty, facilitates better in-distribution performance. Further, while generating diverse synthetic data does not as strongly benefit in-distribution performance, filtering for more diverse data facilitates more robust OOD generalization. We also confirm the existence of model and data scaling laws for synthetically generated problems, which positively benefit downstream model generalization.

推理生成合成数据质量多样性数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。