arXiv:2510.24616stat.MLcond-mat.dis-nn2025-10被引 18

用统计物理方法研究深度网络在过拟合区的最优学习规律。

Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation

  • 构建宽度与输入维度成比例的多层感知机,实现特征学习。
  • 在参数量与数据量相当的过拟合区,模型性能随数据增加分阶段提升。
  • 揭示了深层网络学习存在层级不均的“专业化”现象,适合研究深度神经网络机制的人参考。

四十年来,统计物理为神经网络分析提供了理论框架,但长期未能解决深度学习中复杂特征学习能力的建模问题,此前研究多局限于窄网络或核方法。本文通过监督学习多层感知机,正面回应该问题:(i) 网络宽度与输入维度同阶,比超宽网络更易产生特征学习,又比窄网络更具表达力;(ii) 聚焦于参数量与数据量相近的挑战性过拟合区域,迫使模型适应任务。采用匹配的教师-学生设定,揭示了学习随机深层目标的根本极限,并识别出随着数据量增加,最优网络所学内容的充分统计量。丰富的学习相变现象涌现:当数据足够时,模型通过向目标“专业化”实现最优性能,但训练算法常陷入理论预测的次优解。专业化从浅层向深层逐层传播,且层内各神经元表现不均。此外,深层目标更难学习。尽管模型结构简单,贝叶斯最优设置仍为理解深度、非线性及有限(比例)宽度如何影响特征学习提供深刻洞见,可能适用于更广泛场景。

原文摘要 · Abstract (English)

For four decades statistical physics has been providing a framework to analyse neural networks. A long-standing question remained on its capacity to tackle deep learning models capturing rich feature learning effects, thus going beyond the narrow networks or kernel methods analysed until now. We positively answer through the study of the supervised learning of a multi-layer perceptron. Importantly, (i) its width scales as the input dimension, making it more prone to feature learning than ultra wide networks, and more expressive than narrow ones or ones with fixed embedding layers; and (ii) we focus on the challenging interpolation regime where the number of trainable parameters and data are comparable, which forces the model to adapt to the task. We consider the matched teacher-student setting. Therefore, we provide the fundamental limits of learning random deep neural network targets and identify the sufficient statistics describing what is learnt by an optimally trained network as the data budget increases. A rich phenomenology emerges with various learning transitions. With enough data, optimal performance is attained through the model's "specialisation" towards the target, but it can be hard to reach for training algorithms which get attracted by sub-optimal solutions predicted by the theory. Specialisation occurs inhomogeneously across layers, propagating from shallow towards deep ones, but also across neurons in each layer. Furthermore, deeper targets are harder to learn. Despite its simplicity, the Bayes-optimal setting provides insights on how the depth, non-linearity and finite (proportional) width influence neural networks in the feature learning regime that are potentially relevant in much more general settings.

深度学习统计物理特征学习过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。