揭示噪声数据下深度模型的分阶段学习机制与过拟合良性化现象
Deep Exploration of Epoch-wise Double Descent in Noisy Data: Signal Separation, Large Activation, and Benign Overfitting
- 通过分解损失曲线,分析模型在噪声数据中的信号演化路径
- 训练后期仍能实现再泛化,验证了良性过拟合的存在
- 发现浅层出现巨大激活,且其强度与输入相关而非输出
深度双下降是解释深度学习泛化能力的关键现象。本研究通过关注内部结构演化,实证考察了在噪声数据下的周期性双下降现象。使用三种不同规模的全连接神经网络,在含30%标签噪声的CIFAR-10数据集上进行训练。通过将损失曲线分解为来自干净与噪声训练数据的信号贡献,分别分析了内部信号的周期演化。主要发现:第一,在双下降阶段,模型即使完全拟合噪声数据,仍可在测试集上实现强再泛化,对应“良性过拟合”状态;第二,噪声数据的学习滞后于干净数据,随着训练推进,其对应的深层内部激活逐渐分离,使模型仅对噪声数据过拟合;第三,所有模型在浅层均出现单一超大激活,该现象被称为“异常值”、“巨激活”或“超级激活”,其幅度与输入模式相关,但与输出模式无关。这些实证结果直接关联了“深度双下降”、“良性过拟合”与“大激活”三大前沿现象,支持提出理解深度双下降的新范式。
原文摘要 · Abstract (English)
Deep double descent is one of the key phenomena underlying the generalization capability of deep learning models. In this study, epoch-wise double descent, which is delayed generalization following overfitting, was empirically investigated by focusing on the evolution of internal structures. Fully connected neural networks of three different sizes were trained on the CIFAR-10 dataset with 30% label noise. By decomposing the loss curves into signal contributions from clean and noisy training data, the epoch-wise evolutions of internal signals were analyzed separately. Three main findings were obtained from this analysis. First, the model achieved strong re-generalization on test data even after perfectly fitting noisy training data during the double descent phase, corresponding to a "benign overfitting" state. Second, noisy data were learned after clean data, and as learning progressed, their corresponding internal activations became increasingly separated in outer layers; this enabled the model to overfit only noisy data. Third, a single, very large activation emerged in the shallow layer across all models; this phenomenon is referred as "outliers," "massive activa-tions," and "super activations" in recent large language models and evolves with re-generalization. The magnitude of large activation correlated with input patterns but not with output patterns. These empirical findings directly link the recent key phenomena of "deep double descent," "benign overfitting," and "large activation", and support the proposal of a novel scenario for understanding deep double descent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。