哪怕1%合成数据也会导致模型性能崩塌,越大模型越严重。
Strong Model Collapse
- 用随机投影模拟大模型,发现合成数据引发强模型坍塌
- 仅1%合成数据就让训练集增大也无效,性能不再提升
- 大模型加剧坍塌,但超临界规模后可部分缓解
在支撑ChatGPT、Llama等大模型训练的缩放定律框架下,我们研究了监督回归任务中的模型坍塌现象。结果表明,即使训练集中仅有极小比例的合成数据(如1%),也会引发严重的模型性能退化:增大训练集无法提升性能。进一步探究发现,当前主流的大模型扩展策略反而可能加剧该问题。在通过可调尺寸随机投影近似神经网络的简化设定中,理论与实证均显示更大模型会放大模型坍塌。有趣的是,理论分析还指出,在超过插值阈值(对超大数据集而言极高)后,更大模型虽不能完全避免坍塌,但可部分缓解其影响。该结论在语言模型和图像前馈网络上均得到实验验证。
原文摘要 · Abstract (English)
Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the existance of a strong form of the model collapse phenomenon, a critical performance degradation due to synthetic data in the training corpus. Our results show that even the smallest fraction of synthetic data (e.g., as little as 1\% of the total training dataset) can still lead to model collapse: larger and larger training sets do not enhance performance. We further investigate whether increasing model size, an approach aligned with current trends in training large language models, exacerbates or mitigates model collapse. In a simplified regime where neural networks are approximated via random projections of tunable size, we both theoretically and empirically show that larger models can amplify model collapse. Interestingly, our theory also indicates that, beyond the interpolation threshold (which can be extremely high for very large datasets), larger models may mitigate the collapse, although they do not entirely prevent it. Our theoretical findings are empirically verified through experiments on language models and feed-forward neural networks for images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。