从学习理论角度分析模型生成内容回流对性能的损害机制
Language Generation with Replay: A Learning-Theoretic View of Model Collapse
- 引入回放对抗者,模拟模型输出重新进入训练数据
- 证明非均匀生成和极限生成场景下回流会引发性能退化
- 解释数据清洗等实践方法为何有效及何时失效
随着规模定律推动大语言模型(LLM)对数据的需求持续增长,训练数据正逼近公开网络文本的极限。与此同时,广泛使用LLM导致网络上机器生成内容激增,二者共同增加了生成文本回流至未来训练数据的可能性,从而加剧了常被称为‘模型坍缩’的性能下降风险。尽管实践中通过数据清洗、水印、合成数据策略或无视等方式应对,但模型坍缩尚未从学习理论视角被系统研究。本文基于语言生成极限框架,提出回放对抗者,将生成器历史输出加入示例流中进行分析。主要贡献是精细刻画了回放何时从根本上限制生成:对于最强的均匀生成概念,回放无害;但对于较弱的非均匀生成与极限生成概念,回放会引发可证明的性能分离。有趣的是,我们的正向结果呼应了实践中广泛采用的启发式方法,如数据清洗、水印和输出过滤;而分离结果揭示了这些方法在特定情境下的失效边界。
原文摘要 · Abstract (English)
As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed. At the same time, widespread LLM usage increases the volume of machine-generated content on the web; together, these trends raise the likelihood of generated text re-entering future training corpora, increasing the associated risk of performance degradation often called model collapse. In practice, model developers address this concern through data cleaning, watermarking, synthetic-data policies, or, in some cases, blissful ignorance. However, the problem of model collapse in generative models has not been examined from a learning-theoretic perspective: we study it through the theoretical lens of the language generation in the limit framework, introducing a replay adversary that augments the example stream with the generator's own past outputs. Our main contribution is a fine-grained learning-theoretic characterization of when replay fundamentally limits generation: while replay is benign for the strongest notion of uniform generation, it provably creates separations for the weaker notions of non-uniform generation and generation in the limit. Interestingly, our positive results mirror heuristics widely used in practice, such as data cleaning, watermarking, and output filtering, while our separations show when these ideas can fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。