arXiv:2412.08864cs.CL2024-12NeurIPS被引 3

用知识图谱生成数学推理数据,高效且多样。

GRIP: A Graph-Based Reasoning Instruction Producer

  • 基于种子数据构建知识图谱,利用显性和隐性关系生成指令
  • 从7.5K样本生成210万条问答对,多样性与规模显著提升
  • 适合需要高质量推理数据的AI研究者与模型训练团队

大规模高质量数据对提升大语言模型的推理能力至关重要。随着公开互联网数据日益稀缺,合成数据成为关键方向。然而现有方法常受限于可扩展性不足、样本多样性差及对种子数据过拟合等问题。本文提出GRIP(Graph-based Reasoning Instruction Producer),通过从种子数据中提取高层次概念构建知识图谱,创新性地利用图中显性和隐性关系,实现大规模、多样化的推理指令合成,并采用开源多模态监督保障数据质量。我们在具有挑战性的数学推理领域应用GRIP,从7.5K个初始样本构建出包含210万条合成问答对的GRIP-MATH数据集。相比同类合成方法,GRIP在可扩展性、多样性上更优,成本显著降低。在数学推理基准测试中,使用GRIP-MATH训练的模型相较基线模型有显著提升,优于此前所有数据合成方法。

原文摘要 · Abstract (English)

Large-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer from limited scalability, insufficient sample diversity, and a tendency to overfit to seed data, which constrains their practical utility. In this paper, we present \textit{\textbf{GRIP}}, a \textbf{G}raph-based \textbf{R}easoning \textbf{I}nstruction \textbf{P}roducer that efficiently synthesizes high-quality and diverse reasoning instructions. \textit{GRIP} constructs a knowledge graph by extracting high-level concepts from seed data, and uniquely leverages both explicit and implicit relationships within the graph to drive large-scale and diverse instruction data synthesis, while employing open-source multi-model supervision to ensure data quality. We apply \textit{GRIP} to the critical and challenging domain of mathematical reasoning. Starting from a seed set of 7.5K math reasoning samples, we construct \textbf{GRIP-MATH}, a dataset containing 2.1 million synthesized question-answer pairs. Compared to similar synthetic data methods, \textit{GRIP} achieves greater scalability and diversity while also significantly reducing costs. On mathematical reasoning benchmarks, models trained with GRIP-MATH demonstrate substantial improvements over their base models and significantly outperform previous data synthesis methods.

推理生成知识图谱数据合成数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。