arXiv:2605.27078cs.LGcs.AI2026-05被引 2

拆解模型学习的两阶段:表征与读出,解释为何性能突然提升或波动。

Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

论文配图:Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
图 1 · 摘自论文原文
  • 将学习分解为表征构建和读出校准两个过程,分析其相对速度变化。
  • 发现性能突变前读出过拟合,表征学习持续但缓慢,非传统‘懒惰-丰富’理论。
  • 提供诊断工具,区分真实泛化与训练异常导致的假象,适合模型可解释性研究者。

训练损失和准确率是监测深度神经网络泛化能力的标准信号。但两种现象使其复杂化:在‘悟道’(grokking)中,训练损失迅速下降,而测试性能在长时间延迟后才突然提升;在逐轮双下降(epoch-wise double descent)中,训练损失单调下降,测试损失或误差则先升后降。现有解释多局限于特定任务,缺乏适用于真实任务和架构的通用分析框架。本文通过分析表征学习(编码器)与读出校准(分类器)两个竞争过程,结合表示几何、神经正切核和线性探测工具,证明二者在整个训练过程中均活跃,其相对速度的波动导致看似异常的泛化动态。在多种任务与架构上应用该分解,发现‘悟道’前读出存在训练偏差,表征学习虽渐进但从未停止,与‘懒惰-丰富’假说相悖。该框架还提供诊断标志,区分虚假与真实泛化:在已报道的MNIST‘悟道’和逐轮双下降案例中,观察到的延迟或非单调泛化实为非常规训练策略引起的表征退化与读出错位所致。这些结果确立了表征-读出分解作为理解学习动态的顶层框架,并为可解释性研究揭示底层算法。

原文摘要 · Abstract (English)

Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improves abruptly only after a long delay; in epoch-wise double descent, train loss decreases monotonically while test loss or error rises and falls. Existing accounts are often task-specific, and a task-agnostic analysis framework for diagnosing and explaining these phenomena across realistic tasks and architectures is missing. We address this challenge by analyzing two competing processes that underlie learning dynamics: representation learning in the encoder and readout calibration in the final classifier. Using tools from representational geometry, neural tangent kernels, and linear probing, we show that both processes are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent, contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalization is shown to arise from representation degradation and readout misalignment induced by non-standard training recipes. Together, these results establish the representation-readout decomposition as a top-down framework for understanding learning dynamics and revealing underlying algorithms for interpretability research.

学习动态表征学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。