arXiv:2502.04235cs.CL2025-02被引 10

用新方法生成7700亿词的多样化语料,解决大模型训练重复数据问题

Reformulation for Pretraining Data Augmentation

  • 将现有语料按风格和受众重写,生成多样且上下文丰富的变体数据
  • 在130亿参数模型上验证,有效缓解数据重复带来的性能下降
  • 适合追求高效大模型训练与数据增强的研究者使用

尽管大语言模型在各类任务中表现优异,但其持续扩展不仅受限于数据稀缺,也受训练过程中过度重复数据引发的性能退化影响。为突破这一关键瓶颈,我们提出轻量级、可扩展的数据增强方法——大规模风格-受众(MGA)重构法,受合成数据方法启发,系统性地将现有语料重构为多样且富含上下文的变体,以减轻重复带来的负面影响。本文同时引入由此产生的7700亿词规模的MGACorpus数据集。实验验证了该方法在缩放场景(最高达130亿参数)下对抗数据重复和过采样的优越性能。综合分析揭示了提示工程对生成质量的影响,并指出标准损失指标评估模型能力时存在的细微偏差。结果表明,MGA为大幅扩充训练数据提供了可靠路径,有效缓解重复瓶颈,推动大模型更高效扩展。

原文摘要 · Abstract (English)

Despite the impressive capabilities of large language models across various tasks, their continued scaling is severely hampered not only by data scarcity but also by the performance degradation associated with excessive data repetition during training. To overcome this critical bottleneck, we propose the Massive Genre-Audience(MGA) reformulation method, a lightweight and scalable data augmentation technique inspired by synthetic data methodologies. MGA systematically reformulates existing corpora into diverse, contextually-rich variations to mitigate the negative effects of repetition, and we introduce this approach along with the resulting 770 billion token MGACorpus in this work. We experimentally validate its core benefit by demonstrating superior performance against data repetition and upsampling in scaling scenarios (up to 13B parameters). Furthermore, comprehensive analysis investigates the role of prompt engineering in generation quality and reveals nuances in evaluating model capabilities using standard loss metrics. Our work shows that MGA provides a reliable pathway to substantially augment training datasets, effectively alleviating repetition bottlenecks and enabling more efficient scaling of large language models.

数据增强大模型训练语料重构降重复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。