arXiv:2605.18022cs.LGcs.AI2026-05

大模型在噪声标签下仍能准确推理,因内部存在可提取的泛化结构。

Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise

论文配图:Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise
图 1 · 摘自论文原文
  • 通过频域方法从噪声模型中提取泛化能力
  • 80%标签噪声下仍达近似完美测试准确率
  • 泛化结构分布于多神经元,非集中于特定子网络

高度过参数化的模型能在存在噪声标签的同时实现良好泛化,但其共存机制尚不明确。本文以带强标签噪声的模运算任务为案例,研究两层神经网络中的行为。实验发现,在适当优化与配置下,更大模型泛化能力更强,而噪声标签被更快记忆;过参数化模型内部形成泛化结构,但输出被拟合噪声标签的需求所抑制。令人惊讶的是,即使标签噪声高达80%,通过频域方法提取内部结构仍可实现接近完美的测试准确率。我们进一步提出一种无需任务依赖的网络拆分方法,将网络划分为泛化与记忆组件,尽管该子网络提升泛化性能,但效果仍不及频域提取法,表明泛化结构分布于多个神经元,亟需新工具来挖掘过参数化网络中的可泛化知识。

原文摘要 · Abstract (English)

Highly over-parameterized models can simultaneously memorize noisy labels and generalize well, yet how these behaviors coexist remains poorly understood. In this work, we investigate the underlying mechanisms of this coexistence using modular arithmetic tasks under heavy label noise. Through extensive experiments on two-layer neural networks, we find that larger models tend to generalize better under appropriate optimization and model configurations, while noisy labels are memorized faster than clean data. Over-parameterized models internally form a generalization structure, but its expression in the output is suppressed by the need to fit noisy labels. Remarkably, even with 80\% label noise, near-perfect test accuracy can be achieved by extracting this internal structure using frequency-based methods. We further propose a task-agnostic method to partition networks into generalization and memorization components. Although this subnetwork improves generalization, it is limited compared with frequency-based extraction, indicating that the generalization structure is distributed across neurons and motivating the development of new tools to retrieve generalizable knowledge from over-parameterized networks.

泛化能力噪声标签频域分析过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。