揭示浅层网络在教师-学生框架下的泛化能力与特征对齐机制
Generalization performance of narrow one-hidden layer networks in the teacher-student setting
- 基于统计物理方法建立宽网络泛化理论框架,仅需少量序参量描述性能
- 发现样本量足够大时,隐层神经元会自发对齐教师特征,进入专业化阶段
- 理论可精准预测不同优化算法在回归分类任务中的泛化误差
理解神经网络在简单输入-输出分布上的泛化特性,是解释其在真实数据集上表现的关键。经典的教师-学生设置——即用教师模型生成数据进行训练——提供了标准的理论测试平台。在此背景下,对具有通用激活函数的全连接单隐层网络的完整理论刻画仍不完善。本文针对大宽度但远小于输入维度的此类网络,发展了一套通用理论框架。借助统计物理方法,我们推导出有限温度(贝叶斯)与经验风险最小化估计器的典型性能的闭式表达式,仅依赖少数序参量。我们发现,当样本数足够大且与网络参数数量成比例时,系统会经历相变,进入专业化相,此时隐层神经元将对齐教师特征。该理论能准确预测使用噪声全批量梯度下降(朗之万动力学)或确定性全批量梯度下降训练的网络在回归和分类任务中的泛化误差。
原文摘要 · Abstract (English)
Understanding the generalization properties of neural networks on simple input-output distributions is key to explaining their performance on real datasets. The classical teacher-student setting, where a network is trained on data generated by a teacher model, provides a canonical theoretical test bed. In this context, a complete theoretical characterization of fully connected one-hidden-layer networks with generic activation functions remains missing. In this work, we develop a general framework for such networks with large width, yet much smaller than the input dimension. Using methods from statistical physics, we derive closed-form expressions for the typical performance of both finite-temperature (Bayesian) and empirical risk minimization estimators in terms of a small number of order parameters. We uncover a transition to a specialization phase, where hidden neurons align with teacher features once the number of samples becomes sufficiently large and proportional to the number of network parameters. Our theory accurately predicts the generalization error of networks trained on regression and classification tasks using either noisy full-batch gradient descent (Langevin dynamics) or deterministic full-batch gradient descent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。