揭示在线Softmax分类中学习曲线呈1/3幂律的边界层机制。
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
- 通过中心化变量建模师生对齐与残差方差演化。
- 预测测试损失和泛化误差均随训练时间呈α⁻¹/³衰减。
- 适用于理解复杂数据结构下的渐近学习行为。
硬标签分类通常使用平滑代理损失进行训练,最典型的是Softmax交叉熵。我们识别出一种渐近机制,即平滑代理损失与离散标签之间的不匹配导致在线教师-学生模型中出现幂律学习曲线。在减去平均logit后,热力学极限下的动力学在中心化变量中闭合:不断增长的师生对齐度D与残差学生方差Δ。在后期,远离教师决策边界的样本已能被自信分类,贡献指数级微小;仅有宽度为O(D⁻¹)的边界层仍活跃,而固定学习率在线梯度下降的噪声维持了非零Δ。作为训练时间α的函数,晚期解给出α⁻¹/³幂律,不仅适用于测试损失,也适用于泛化误差ε_g(即1减测试准确率)。该速率远慢于同模型下的贝叶斯最优参考速率α⁻¹。我们进一步表明,学习率调度可使泛化误差趋近α⁻¹/²幂律。模拟结果支持预测的序参数动态和学习曲线。在相关高斯输入与白化预训练特征上的受控实验显示,数据结构可主导瞬态行为。因此,本结果为渐近、互补机制,而非神经缩放定律谱解释的替代方案。
原文摘要 · Abstract (English)
Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between smooth surrogate and discrete labels produces power-law learning curves in an online teacher-student model. After subtracting the mean logit, the thermodynamic-limit dynamics close in centered variables: a growing centered student-teacher alignment $D$ and the residual student variance $Δ$. At late times, examples away from teacher decision boundaries are already classified confidently and contribute exponentially little. Only boundary layers of width $O(D^{-1})$ remain active, while the noise of fixed-learning-rate online gradient descent maintains a nonzero $Δ$. As a function of the training time $α$ the late-time solution yields a $α^{-1/3}$ power law not only for the test loss but also for the generalization error $ε_g$, i.e., one minus test accuracy. This is much slower than the $α^{-1}$ Bayes-optimal reference for the same model. We further show that learning-rate schedules can improve the generalization error towards a $ε_g \sim α^{-1/2}$ power law. Simulations support the predicted order parameter dynamics and learning curves. Controlled experiments with correlated Gaussian inputs and whitened pretrained features show that data structure can dominate transients. Therefore, our result is an asymptotic, complementary mechanism rather than an alternative to spectral explanations of neural scaling laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。