用技能图谱精选数学数据,让大模型预训练更高效、更精准。
MASS: Mathematical Data Selection via Skill Graphs for Pretraining Large Language Models
- 构建数学技能图谱,量化数据与推理能力的关联。
- 减少50%-70%训练词数,性能持平原数据集。
- 适合数学推理类大模型预训练,提升训练效率。
高质量数据在大语言模型(LLMs)的预训练和微调中至关重要,甚至决定其性能上限。尽管已有众多数据选择方法,但多数聚焦通用场景,忽视领域特异性。本文提出MASS框架,基于技能图谱实现数学推理领域的大模型预训练数据精选。通过分析参考数据集中的数学技能及其相互关系,构建技能图谱,指导目标数据的质量评分,从而选出高价值子集用于预训练。实验表明,无论模型规模(1B/7B)或数据来源(网络数据/合成数据),MASS均表现优异:在效率方面,使用MASS选中的子集训练的模型仅需原数据50%-70%的训练词数即可达到相近性能;在有效性方面,在相同训练词数下,模型性能提升3.3%-5.9%。结果验证了MASS在提升预训练效率与效果方面的潜力。
原文摘要 · Abstract (English)
High-quality data plays a critical role in the pretraining and fine-tuning of large language models (LLMs), even determining their performance ceiling to some degree. Consequently, numerous data selection methods have been proposed to identify subsets of data that can effectively and efficiently enhance model performance. However, most of these methods focus on general data selection and tend to overlook the specific nuances of domain-related data. In this paper, we introduce MASS, a \textbf{MA}thematical data \textbf{S}election framework using the \textbf{S}kill graph for pretraining LLMs in the mathematical reasoning domain. By taking into account the unique characteristics of mathematics and reasoning, we construct a skill graph that captures the mathematical skills and their interrelations from a reference dataset. This skill graph guides us in assigning quality scores to the target dataset, enabling us to select the top-ranked subset which is further used to pretrain LLMs. Experimental results demonstrate the efficiency and effectiveness of MASS across different model sizes (1B and 7B) and pretraining datasets (web data and synthetic data). Specifically, in terms of efficiency, models trained on subsets selected by MASS can achieve similar performance to models trained on the original datasets, with a significant reduction in the number of trained tokens - ranging from 50\% to 70\% fewer tokens. In terms of effectiveness, when trained on the same amount of tokens, models trained on the data selected by MASS outperform those trained on the original datasets by 3.3\% to 5.9\%. These results underscore the potential of MASS to improve both the efficiency and effectiveness of pretraining LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。