arXiv:2502.08924cs.LGcs.AI2025-02NeurIPS被引 17

弱数据也能训练出强大模型,关键在动态聚焦难点样本

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

  • 借鉴提升算法思想,用弱数据动态聚焦难例进行训练
  • 实验验证:持续优化性能,避免模型训练陷入停滞
  • 适合关注合成数据训练、模型泛化能力的研究者

合成数据在大语言模型训练中的作用日益重要。然而,研究表明,若缺乏有效筛选,合成数据会导致模型性能在多次训练迭代后停滞甚至“崩溃”。本文提出一个理论框架,分析训练大模型所需的数据筛选程度,以确保性能持续提升。该分析受经典提升算法启发,其核心思想是利用非常弱的学习器构建任意高性能分类器。所分析的方法涵盖了近期多种基于合成数据训练大模型的技术,揭示了它们成功的原因,并为未来改进提供思路。实验结果验证了理论假设,表明将标注资源动态集中在最具挑战性的样本上——如同提升算法集中弱学习器的注意力——可显著提升模型表现。

原文摘要 · Abstract (English)

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even "collapse", after many training iterations. In this paper, we formalize this question and develop a theoretical framework to investigate how much curation is needed in order to ensure that LLM performance continually improves. Our analysis is inspired by boosting, a classic machine learning technique that leverages a very weak learning algorithm to produce an arbitrarily good classifier. The approach we analyze subsumes many recently proposed methods for training LLMs on synthetic data, and thus our analysis sheds light on why they are successful, and also suggests opportunities for future improvement. We present experiments that validate our theory, and show that dynamically focusing labeling resources on the most challenging examples -- in much the same way that boosting focuses the efforts of the weak learner -- leads to improved performance.

大模型训练合成数据提升算法模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。