从统计视角重新理解神经网络的特征学习机制,突破传统核方法局限。
Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories
- 将神经网络视为自适应特征模型,超越固定核理论
- 提出过参数化高斯序列模型作为特征学习研究原型
- 为未来分析神经网络泛化能力提供新思路,适合理论研究者
神经网络的核心优势在于其特征学习能力,但训练动态复杂,理论分析困难。本文从统计角度审视特征学习及其对泛化的影响。回顾了神经正切核(NTK)理论及近期核回归成果,这些理论解释了充分宽网络的泛化性能。然而,固定核理论存在局限,本文进一步探讨特征学习的最新理论进展。不再局限于固定特征,转而将神经网络视为自适应特征模型。最后,提出一个过参数化高斯序列模型作为自适应特征模型的原型,用于研究特征学习特性,并为未来神经网络的理论分析提供动机。
原文摘要 · Abstract (English)
A primary advantage of neural networks lies in their feature learning characteristics, which is challenging to theoretically analyze due to the complexity of their training dynamics. We examine feature learning and its potential benefits for generalization from a statistical perspective. After reviewing the neural tangent kernel (NTK) theory and recent results in kernel regression, which address the generalization issue of sufficiently wide neural networks, we examine limitations and implications of the fixed kernel theory (as the NTK theory) and review recent theoretical advancements in feature learning. Moving beyond theories with fixed features, we consider neural networks as adaptive feature models. Finally, we propose an over-parameterized Gaussian sequence model as a prototype for the adaptive feature model to study feature learning characteristics and motivate their future analysis for neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。