arXiv:2409.15376cs.LGcs.AI2024-09EMNLP被引 9

用可控生成法扩充数学题库,提升模型泛化能力

ControlMath: Controllable Data Generation Promotes Math Generalist Models

  • 通过方程生成与双代理迭代,生成多样化数学应用题
  • 构建19万道题的ControlMathQA数据集,显著提升模型泛化性能
  • 适合需要跨领域数学推理的模型训练与评估

利用大语言模型进行数据增强在数学推理任务中已取得令人鼓舞的结果。然而,这些方法在问题多样性方面存在局限,可能仅限于特定领域或分布的数据生成。为此,我们提出ControlMath,一种包含方程生成模块和两个基于LLM的代理的迭代方法。该模块生成多样化的方程,由Problem-Crafter代理将其转化为数学应用题,Reverse-Agent则过滤并筛选高质量数据,遵循“少即是多”原则,在更少数据点下获得更好效果。该方法可生成不限于特定领域或分布的多样化数学问题。最终,我们构建了包含19万道数学应用题的ControlMathQA数据集。大量实验表明,将该数据集与GSM8K等领域内数据集结合,能有效提升模型的数学推理泛化能力,使其在域内及域外任务中表现均有所提升。

原文摘要 · Abstract (English)

Utilizing large language models (LLMs) for data augmentation has yielded encouraging results in mathematical reasoning. However, these approaches face constraints in problem diversity, potentially restricting them to in-domain/distribution data generation. To this end, we propose ControlMath, an iterative method involving an equation-generator module and two LLM-based agents. The module creates diverse equations, which the Problem-Crafter agent then transforms into math word problems. The Reverse-Agent filters and selects high-quality data, adhering to the "less is more" principle, achieving better results with fewer data points. This approach enables the generation of diverse math problems, not limited to specific domains or distributions. As a result, we collect ControlMathQA, which involves 190k math word problems. Extensive results prove that combining our dataset with in-domain datasets like GSM8K can help improve the model's mathematical ability to generalize, leading to improved performances both within and beyond specific domains.

数学推理数据生成大模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。