arXiv:2412.14872cs.CL2024-12被引 4

证明自回归语言模型在有限真实数据下必然崩溃

Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set

  • 理论上证明:一旦数据集混入生成数据且无新真实数据,模型终将崩溃
  • 无论生成数据量多小,只要持续循环训练,崩溃不可避免
  • 提示:解决崩溃关键在生成数据质量,而非限制数量

自回归语言模型(LMs)被广泛用于数据稀缺领域生成数据以训练新模型,弥补真实世界数据的不足。先前研究曾实验发现,当模型在递归生成的数据上训练时会发生崩溃。本文首次提出理论证明:一旦语料库(如万维网子集)开始包含生成数据,且不再引入新的真实世界数据,则无论每个模型生成的数据量多么微小,只要经过足够时间,模型崩溃必然发生。该结论表明,通过限制合成数据量来缓解崩溃的方法从根本上是无效的。因此,避免崩溃的关键在于确保合成数据的质量。

原文摘要 · Abstract (English)

Auto-regressive language models (LMs) have been widely used to generate data in data-scarce domains to train new LMs, compensating for the scarcity of real-world data. Previous work experimentally found that LMs collapse when trained on recursively generated data. This paper presents a theoretical proof: once a corpus (such as a subset of the World Wide Web) begins to incorporate generated data and no new real-world data is added to the corpus, then no matter how small the amount of data each LM generates and contributes to the corpus, LM collapse is inevitable after sufficient time. This finding suggests that attempts to mitigate collapse by limiting the quantity of synthetic data in the corpus are fundamentally insufficient. Instead, avoiding collapse hinges on ensuring the quality of synthetic data.

语言模型模型崩溃生成数据理论证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。