通过梯度谱特征揭示噪声标签下的记忆现象
Fisher Rank Inflation: A Spectral Signature of Memorization under Label Noise

- 分析最后一层梯度的中心散度谱,发现记忆阶段有效秩会短暂上升
- 噪声标签使低能方向注入谱质量,峰值有效秩达97.09(60%噪声时)
- 适用于检测模型何时开始记忆错误标签,适合研究模型鲁棒性者
在标签噪声下训练的深度网络会先学习干净结构,再记忆错误标签。我们发现这一转变会在逐样本最后层梯度的中心散度中留下谱特征:记忆阶段有效秩短暂膨胀,拟合后又收缩。此现象称为Fisher秩膨胀。噪声标签通过向低能量或未使用特征方向注入谱质量,提升梯度谱熵。我们推导了一阶留一法归因公式,识别出错误样本贡献强于正确样本的条件,并解释了归因信号在标准化Fisher-梯度谱稳定后减弱的原因。在CIFAR-10、CIFAR-100和CIFAR-10N上测试,使用SmallCNN、ResNet18和Vision Transformer,所有设置下均观察到一致的秩膨胀-坍缩轨迹。峰值秩检查点处,最高秩贡献样本中噪声比例达69.2%至96.2%(五次种子合成噪声实验),在CIFAR-10N上为94.4%±1.9%。一阶谱归因与精确留一法贡献高度一致,卷积模型中尤为接近,视觉变换器中仍保持富集。峰值有效秩随噪声严重性单调上升,从清洁训练的28.88±1.95增至60%噪声时的97.09±1.78。在若干设置中,回顾性识别的秩膨胀起始时间早于可见的测试性能下降。这些结果确立了Fisher秩膨胀作为连接错误样本富集、噪声严重性与结构学习到记忆转变的谱签名。
原文摘要 · Abstract (English)
Deep networks trained with label noise often learn clean structure before memorizing corrupted labels. We show that this transition leaves a spectral signature in the centered scatter of per-example last-layer gradients. Its effective rank transiently expands during memorization and contracts after corrupted labels are fit. We call this phenomenon Fisher Rank Inflation. Corrupted labels increase effective rank by injecting spectral mass into low-energy or previously unused eigendirections, increasing the entropy of the gradient spectrum. We derive a first-order leave-one-out attribution formula, identify conditions under which corrupted examples contribute more strongly than clean examples, and explain why attribution signals weaken once the normalized Fisher-gradient spectrum stabilizes. We test these predictions on CIFAR-10, CIFAR-100, and CIFAR-10N using SmallCNN, ResNet18, and Vision Transformers. Across settings, Fisher effective rank exhibits a consistent inflation--collapse trajectory aligned with memorization. At peak-rank checkpoints, corrupted examples are enriched among the highest rank-contributing samples, with top-100 noisy fractions from \(69.2\%\) to \(96.2\%\) across five-seed synthetic-corruption experiments and \(94.4\%\pm1.9\%\) on CIFAR-10N. First-order spectral attribution closely matches exact leave-one-out contributions in convolutional models and remains enriched in the Vision Transformer. Peak effective rank increases monotonically with corruption severity, from \(28.88\pm1.95\) under clean training to \(97.09\pm1.78\) at \(60\%\) corruption. In several settings, the retrospectively identified onset of rank inflation precedes observable test degradation. These results establish Fisher Rank Inflation as a spectral signature connecting corrupted-example enrichment, corruption severity, and the transition from structure learning to memorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。