arXiv:2501.18845cs.CL2025-01综述被引 42

综述大模型文本数据增强方法,助小数据训练更高效

Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

  • 按生成方式分四类:简单、提示词、检索增强与混合方法
  • 引入外部知识提升生成内容真实性,减少虚假信息
  • 适合想提升小样本训练效果的研究者和工程师

预训练语言模型规模日益增大,在诸多应用中表现优异,但通常需要大量训练数据以充分训练。数据不足易导致模型过拟合,难以应对复杂任务。基于大规模语料训练的大语言模型具备强大的文本生成能力,可有效提升数据质量和数量,对数据增强至关重要。具体而言,通过个性化任务的提示模板引导大模型生成所需内容;近期基于检索的技术进一步提升了模型在数据增强中的表达能力,通过引入外部知识使生成内容更具事实依据。本文系统梳理大模型数据增强技术,将其分为简单增强、提示词增强、检索增强与混合增强四类;总结后处理方法,显著提升生成数据质量并过滤不真实内容;归纳常用任务与评估指标;最后分析现有挑战与未来机遇,为数据增强研究提供方向。

原文摘要 · Abstract (English)

The increasing size and complexity of pre-trained language models have demonstrated superior performance in many applications, but they usually require large training datasets to be adequately trained. Insufficient training sets could unexpectedly make the model overfit and fail to cope with complex tasks. Large language models (LLMs) trained on extensive corpora have prominent text generation capabilities, which improve the quality and quantity of data and play a crucial role in data augmentation. Specifically, distinctive prompt templates are given in personalised tasks to guide LLMs in generating the required content. Recent promising retrieval-based techniques further improve the expressive performance of LLMs in data augmentation by introducing external knowledge to enable them to produce more grounded-truth data. This survey provides an in-depth analysis of data augmentation in LLMs, classifying the techniques into Simple Augmentation, Prompt-based Augmentation, Retrieval-based Augmentation and Hybrid Augmentation. We summarise the post-processing approaches in data augmentation, which contributes significantly to refining the augmented data and enabling the model to filter out unfaithful content. Then, we provide the common tasks and evaluation metrics. Finally, we introduce existing challenges and future opportunities that could bring further improvement to data augmentation.

大模型数据增强提示工程文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。