arXiv:2505.24849stat.MLcond-mat.dis-nn2025-05被引 18

研究大宽度贝叶斯神经网络在拟合点附近的统计力学行为,揭示特征学习的多阶段转变。

Statistical mechanics of extensive-width Bayesian neural networks near interpolation

  • 采用两层全连接网络,在隐藏层与输入维数同比例增大的条件下分析贝叶斯最优学习。
  • 数据量增加时出现多种学习相变,教师特征贡献越强,所需数据越少即可学习。
  • 数据稀缺时模型仅学非线性组合权重,需足够数据才实现权重对齐,存在训练算法难以突破的统计-计算间隙。

过去三十年,统计物理为神经网络分析提供了理论框架,但可解析的模型(如感知机、随机特征模型、核机器)仍远简单于实际应用中的网络。本文通过统计物理方法分析一个具有通用权重分布和激活函数的两层全连接网络的监督学习问题,其隐藏层虽大但与输入维度同比例增长,比无限宽网络更真实且更具表达力。研究聚焦教师-学生场景下的贝叶斯最优学习,即数据由同架构网络生成。在参数量与数据量相当的拟合区域附近,特征学习开始显现。分析揭示丰富现象学:随着数据量增加,出现多种学习相变;教师特征对输出贡献越强,所需学习数据越少。数据稀缺时,模型仅学习教师权重的非线性组合,未发生权重对齐;只有当数据充足时,权重对齐才可能出现,但该过程可能因统计-计算间隙而难以被实际训练算法发现。

原文摘要 · Abstract (English)

For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models and committee machines with few neurons, remained simple compared to those used in applications. In this paper we help reducing the gap between practical networks and their theoretical understanding through a statistical physics analysis of the supervised learning of a two-layer fully connected network with generic weight distribution and activation function, whose hidden layer is large but remains proportional to the inputs dimension. This makes it more realistic than infinitely wide networks where no feature learning occurs, but also more expressive than narrow ones or with fixed inner weights. We focus on the Bayes-optimal learning in the teacher-student scenario, i.e., with a dataset generated by another network with the same architecture. We operate around interpolation, where the number of trainable parameters and of data are comparable and feature learning emerges. Our analysis uncovers a rich phenomenology with various learning transitions as the number of data increases. In particular, the more strongly the features (i.e., hidden neurons of the target) contribute to the observed responses, the less data is needed to learn them. Moreover, when the data is scarce, the model only learns non-linear combinations of the teacher weights, rather than "specialising" by aligning its weights with the teacher's. Specialisation occurs only when enough data becomes available, but it can be hard to find for practical training algorithms, possibly due to statistical-to-computational~gaps.

神经网络统计力学贝叶斯学习特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。