用知识图谱生成数学推理数据,高效且多样。
GRIP: A Graph-Based Reasoning Instruction Producer
- 基于种子数据构建知识图谱,利用显性和隐性关系生成指令
- 从7.5K样本生成210万条问答对,多样性与规模显著提升
- 适合需要高质量推理数据的AI研究者与模型训练团队
大规模高质量数据对提升大语言模型的推理能力至关重要。随着公开互联网数据日益稀缺,合成数据成为关键方向。然而现有方法常受限于可扩展性不足、样本多样性差及对种子数据过拟合等问题。本文提出GRIP(Graph-based Reasoning Instruction Producer),通过从种子数据中提取高层次概念构建知识图谱,创新性地利用图中显性和隐性关系,实现大规模、多样化的推理指令合成,并采用开源多模态监督保障数据质量。我们在具有挑战性的数学推理领域应用GRIP,从7.5K个初始样本构建出包含210万条合成问答对的GRIP-MATH数据集。相比同类合成方法,GRIP在可扩展性、多样性上更优,成本显著降低。在数学推理基准测试中,使用GRIP-MATH训练的模型相较基线模型有显著提升,优于此前所有数据合成方法。
原文摘要 · Abstract (English)
Large-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer from limited scalability, insufficient sample diversity, and a tendency to overfit to seed data, which constrains their practical utility. In this paper, we present \textit{\textbf{GRIP}}, a \textbf{G}raph-based \textbf{R}easoning \textbf{I}nstruction \textbf{P}roducer that efficiently synthesizes high-quality and diverse reasoning instructions. \textit{GRIP} constructs a knowledge graph by extracting high-level concepts from seed data, and uniquely leverages both explicit and implicit relationships within the graph to drive large-scale and diverse instruction data synthesis, while employing open-source multi-model supervision to ensure data quality. We apply \textit{GRIP} to the critical and challenging domain of mathematical reasoning. Starting from a seed set of 7.5K math reasoning samples, we construct \textbf{GRIP-MATH}, a dataset containing 2.1 million synthesized question-answer pairs. Compared to similar synthetic data methods, \textit{GRIP} achieves greater scalability and diversity while also significantly reducing costs. On mathematical reasoning benchmarks, models trained with GRIP-MATH demonstrate substantial improvements over their base models and significantly outperform previous data synthesis methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。