arXiv:2603.02293cs.LGcs.AI2026-03

发现过参数模型在噪声下会分离信号与噪声,导致泛化失效。

The Malignant Tail: Spectral Segregation of Label Noise in Over-Parameterized Networks

  • 通过频谱探针揭示梯度下降使噪声进入高频正交空间
  • 训练后可通过低秩截断恢复最优泛化能力
  • 适合研究模型过拟合与鲁棒性提升的学者

尽管隐式正则化在低噪声下可实现良性过拟合,但近期理论预测当噪声与信号比升高时会出现有害过拟合的突变。我们实验分离出这一转变的几何机制:恶性尾部(Malignant Tail),即网络将信号与噪声功能上分离——系统性语义特征被压缩至低秩子空间,而随机标签噪声则被推向高频正交分量,区别于与数据扭曲对齐的噪声。通过训练过程的频谱线性探针分析,我们发现梯度下降(SGD)无法抑制此类噪声,反而将其隐式偏导向高频正交子空间,维持了信号-噪声的可分离性。该几何分离不同于未训练模型中的方差降低。在已训练网络中,SGD主动实现噪声分离,使得事后显式频谱截断(d << D)可精准剪除噪声主导子空间,恢复收敛模型中潜藏的最佳泛化性能。相比不稳定的早期停止,几何截断提供稳定的事后干预。结果表明,在标签噪声下,多余的频谱容量并非无害冗余,而是允许噪声记忆的潜在结构缺陷,需通过显式秩约束过滤随机扰动以实现稳健泛化。

原文摘要 · Abstract (English)

While implicit regularization facilitates benign overfitting in low-noise regimes, recent theoretical work predicts a sharp phase transition to harmful overfitting as the noise-to-signal ratio increases. We experimentally isolate the geometric mechanism of this transition: the Malignant Tail, a failure mode where networks functionally segregate signal and noise, reducing coherent semantic features into low-rank subspaces while pushing stochastic label noise into high-frequency orthogonal components, distinct from systematic or corruption-aligned noise. Through a Spectral Linear Probe of training dynamics, we demonstrate that Stochastic Gradient Descent (SGD) fails to suppress this noise, instead implicitly biasing it toward high-frequency orthogonal subspaces, effectively preserving signal-noise separability. We show that this geometric separation is distinct from simple variance reduction in untrained models. In trained networks, SGD actively segregates noise, allowing post-hoc Explicit Spectral Truncation (d << D) to surgically prune the noise-dominated subspace. This approach recovers the optimal generalization capability latent in the converged model. Unlike unstable temporal early stopping, Geometric Truncation provides a stable post-hoc intervention. Our findings suggest that under label noise, excess spectral capacity is not harmless redundancy but a latent structural liability that allows for noise memorization, necessitating explicit rank constraints to filter stochastic corruptions for robust generalization.

过拟合标签噪声频谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。