系统梳理数据高效大模型后训练方法,解决标注成本高、数据边际收益递减难题。
A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
- 从数据视角构建五类高效方法分类体系:数据筛选、质量提升、合成生成、压缩蒸馏、自演化生态。
- 提出首个数据驱动的大模型后训练系统综述,覆盖代表性技术与未来方向。
- 适合关注数据利用效率、低资源训练的AI研究者和工程师阅读。
大语言模型(LLMs)的后训练对释放其任务泛化能力和领域特定能力至关重要。然而,当前大模型后训练面临显著的数据挑战,包括人工标注成本高及数据规模增加带来的边际收益递减。因此,实现数据高效的后训练成为关键研究课题。本文首次从数据中心视角系统性地综述了数据高效的大型语言模型后训练方法。我们提出了一个涵盖数据选择、数据质量增强、合成数据生成、数据蒸馏与压缩以及自演化数据生态系统的分类体系。总结了各类别中的代表性方法,并指出了未来的研究方向。通过分析数据高效后训练中的挑战,我们揭示了开放问题并提出了潜在的研究路径。希望本工作能激发进一步探索大规模模型训练中数据利用潜力的研究。
原文摘要 · Abstract (English)
Post-training of Large Language Models (LLMs) is crucial for unlocking their task generalization potential and domain-specific capabilities. However, the current LLM post-training paradigm faces significant data challenges, including the high costs of manual annotation and diminishing marginal returns on data scales. Therefore, achieving data-efficient post-training has become a key research question. In this paper, we present the first systematic survey of data-efficient LLM post-training from a data-centric perspective. We propose a taxonomy of data-efficient LLM post-training methods, covering data selection, data quality enhancement, synthetic data generation, data distillation and compression, and self-evolving data ecosystems. We summarize representative approaches in each category and outline future research directions. By examining the challenges in data-efficient LLM post-training, we highlight open problems and propose potential research avenues. We hope our work inspires further exploration into maximizing the potential of data utilization in large-scale model training. Paper List: https://github.com/luo-junyu/Awesome-Data-Efficient-LLM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。