揭示梯度训练隐式偏差如何催生从感知机到深度网络的学习曲线缩放规律
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
- 基于范数复杂度度量,发现学习过程中的两类动态缩放规律
- 动态规律统一解释了收敛时的测试误差缩放定律,涵盖多种模型与数据集
- 通过感知机理论推导验证,阐明梯度下降隐式偏差是根源
深度学习中的缩放定律——模型性能与资源增长之间的经验幂律关系——已成为跨架构、数据集和任务的显著规律。这些定律对指导前沿模型设计具有重要意义,能量化增加数据或模型规模的收益,并暗示机器学习可解释性的基础。然而,现有研究多关注训练结束时的渐近行为。本文通过分析完整训练动态,揭示了两种新的动态缩放规律,它们描述了性能随不同范数复杂度度量的变化过程。二者结合可恢复已知的收敛测试误差缩放规律。研究在CNN、ResNet和视觉变压器上于MNIST、CIFAR-10和CIFAR-100上均获得一致结果。此外,我们使用单层感知机与逻辑损失训练,通过解析推导验证了新规律,并通过梯度训练诱导的隐式偏差加以解释。
原文摘要 · Abstract (English)
Scaling laws in deep learning -- empirical power-law relationships linking model performance to resource growth -- have emerged as simple yet striking regularities across architectures, datasets, and tasks. These laws are particularly impactful in guiding the design of state-of-the-art models, since they quantify the benefits of increasing data or model size, and hint at the foundations of interpretability in machine learning. However, most studies focus on asymptotic behavior at the end of training. In this work, we describe a richer picture by analyzing the entire training dynamics: we identify two novel \textit{dynamical} scaling laws that govern how performance evolves as function of different norm-based complexity measures. Combined, our new laws recover the well-known scaling for test error at convergence. Our findings are consistent across CNNs, ResNets, and Vision Transformers trained on MNIST, CIFAR-10 and CIFAR-100. Furthermore, we provide analytical support using a single-layer perceptron trained with logistic loss, where we derive the new dynamical scaling laws, and we explain them through the implicit bias induced by gradient-based training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。