arXiv:2410.12881cs.AIcs.CL2024-10ICLR被引 14

用数学对话生成新数据,显著提升大模型推理能力

MIND: Math Informed syNthetic Dialogues for Pretraining LLMs

  • 基于数学问答构建对话数据,模拟知识差距增强多样性
  • 预训练后在GSM8K上提升13.42%,MATH提升2.30%
  • 适合想提升数学与复杂推理能力的研究者使用

合成数据在提升大语言模型预训练质量方面已得到广泛研究,但在需要多跳推理和数学能力的复杂任务中表现不足,因现有合成数据难以补充原始语料的空白。本文提出一种大规模、多样化的数学引导对话生成方法(MIND),基于OpenWebMath(OWM)生成合成对话数据,构建新的数学语料MIND-OWM。实验表明,参与者间存在知识差异对生成高质量数学数据至关重要。通过优化合成数据与原始数据的格式与融合方式,可最大化数学推理性能提升,强调需重构原始数据而非直接使用。相较于仅使用原始数据预训练的模型,基于MIND-OWM训练的模型在数学推理任务中表现显著提升:GSM8K提高13.42%,MATH提高2.30%;在专业知识任务(MMLU: +4.55%,MMLU-STEM: +4.28%)和通用推理任务(GENERAL REASONING: +2.51%)中同样表现出色。

原文摘要 · Abstract (English)

The utility of synthetic data to enhance pretraining data quality and hence to improve downstream task accuracy has been widely explored in recent large language models (LLMs). Yet, these approaches fall inadequate in complex, multi-hop and mathematical reasoning tasks as the synthetic data typically fails to add complementary knowledge to the existing raw corpus. In this work, we propose a novel large-scale and diverse Math Informed syNthetic Dialogue (MIND) generation method that improves the mathematical reasoning ability of LLMs. Specifically, using MIND, we generate synthetic conversations based on OpenWebMath (OWM), resulting in a new math corpus, MIND-OWM. Our experiments with different conversational settings reveal that incorporating knowledge gaps between dialog participants is essential for generating high-quality math data. We further identify an effective way to format and integrate synthetic and raw data during pretraining to maximize the gain in mathematical reasoning, emphasizing the need to restructure raw data rather than use it as-is. Compared to pretraining just on raw data, a model pretrained on MIND-OWM shows significant boost in mathematical reasoning (GSM8K: +13.42%, MATH: +2.30%), including superior performance in specialized knowledge (MMLU: +4.55%, MMLU-STEM: +4.28%) and general purpose reasoning tasks (GENERAL REASONING: +2.51%).

数学推理合成数据对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。