arXiv:2410.12896cs.CL2024-10综述被引 52

综述大模型数据生成技术,破解高质量数据短缺难题

A Survey on Data Synthesis and Augmentation for Large Language Models

  • 系统梳理大模型全生命周期的数据增强与合成方法
  • 指出当前数据生成面临效率与质量双重瓶颈
  • 适合研究者快速定位合适的数据策略

大语言模型(LLMs)的成功高度依赖于海量、多样且高质量的数据。然而,高质量数据的增长速度远低于训练数据集的扩张,导致数据枯竭危机迫在眉睫。这凸显了提升数据效率和探索新数据源的紧迫性。在此背景下,合成数据成为有前景的解决方案。目前,数据生成主要分为数据增强和数据合成两大类。本文全面回顾并总结了大模型生命周期中各阶段的数据生成技术,涵盖数据准备、预训练、微调、指令调优、偏好对齐及应用。此外,我们分析了现有方法的局限性,并探讨未来发展方向。旨在帮助研究者清晰理解各类方法,快速选择合适的数据生成策略,为后续研究提供参考。

原文摘要 · Abstract (English)

The success of Large Language Models (LLMs) is inherently linked to the availability of vast, diverse, and high-quality data for training and evaluation. However, the growth rate of high-quality data is significantly outpaced by the expansion of training datasets, leading to a looming data exhaustion crisis. This underscores the urgent need to enhance data efficiency and explore new data sources. In this context, synthetic data has emerged as a promising solution. Currently, data generation primarily consists of two major approaches: data augmentation and synthesis. This paper comprehensively reviews and summarizes data generation techniques throughout the lifecycle of LLMs, including data preparation, pre-training, fine-tuning, instruction-tuning, preference alignment, and applications. Furthermore, We discuss the current constraints faced by these methods and investigate potential pathways for future development and research. Our aspiration is to equip researchers with a clear understanding of these methodologies, enabling them to swiftly identify appropriate data generation strategies in the construction of LLMs, while providing valuable insights for future exploration.

数据合成大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。