用文化演化理论解释大模型自训练时的性能退化,发现组合性先升后降。
Model Collapse as Cultural Evolution
- 基于文化演化中的迭代学习理论,提出可验证的五个预测。
- 自训练10代后组合性先增后减,仅任务相关过滤能维持其稳定。
- 首次在大模型规模上验证压缩与沟通的权衡,适合优化自训练流程的研究者。
大模型在自身输出上持续训练导致的模型坍缩现象,虽已被统计描述,但缺乏语言层面的解释——哪些结构退化、按何顺序、为何如此。本文引入文化演化中的迭代学习理论填补空白,推导出五个可检验的预测,区分出唯一判别性预测与验证性预测,并通过在英语、德语和土耳其语中对LLaMA-2-7B和Mistral-7B进行10代自训练实验加以验证。关键判别性发现:组合性在无过滤自训练下呈现非单调轨迹(先上升后下降)。该特征在高度规则的种子数据下仍存在(排除噪声消除影响),且仅任务导向过滤可维持其稳定性,首次在大模型尺度提供压缩-通信权衡的证据。所有预测均获大效应量支持(Hedges' g > 1.6;BF₁₀ > 100),模型正则化梯度与人类行为数据高度一致(R² = 0.94)。研究将模型坍缩重新定义为文化传递现象,为自训练流程设计提供具体原则。
原文摘要 · Abstract (English)
Model collapse, the progressive degradation of LLMs trained on their own outputs, has been characterized statistically but lacks a linguistic explanation for which structures degrade, in what order, and why. We show that iterated learning theory from cultural evolution fills this gap. We derive five falsifiable predictions, distinguish those uniquely discriminative for the theory from confirmatory ones, and test them by self-training LLaMA-2-7B and Mistral-7B over 10 generations in English, German, and Turkish. The critical discriminative finding: compositionality follows a non-monotonic trajectory (initially rising, then falling) under unfiltered self-training. This signature persists with maximally regular seed data (ruling out noise removal) and is sustained only by task-grounded filtering, not random filtering, providing the first LLM-scale evidence for the compression-communication tradeoff. All predictions are confirmed with large effect sizes (Hedges' $g > 1.6$; $\mathrm{BF}_{10} > 100$), and LLM regularization gradients closely match human behavioral data ($R^2 = 0.94$). These results reframe model collapse as a cultural transmission phenomenon and yield concrete principles for self-training pipeline design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。