提出新理论框架,解析高维深度学习的非线性行为
Random Matrix Theory for Deep Learning: Beyond Eigenvalues of Linear Models
- 引入高维等价概念,统一处理高维、非线性与谱函数分析
- 揭示线性、浅层与深层网络在高维下的泛化性能规律
- 为理解双下降、标度律等现象提供统一理论支撑
现代机器学习与深度神经网络常作用于高维数据并依赖过参数化模型,传统低维直觉失效。当数据维度、样本量与模型参数量均大且可比时,会出现新颖甚至反直觉的行为。本文将经典随机矩阵理论(RMT)从线性模型的特征值分析拓展至非线性模型,提出‘高维等价’概念,统一并推广了确定性等价与线性等价,系统解决高维、非线性及通用谱函数分析三大挑战。基于该框架,精确刻画了线性模型、非线性浅层网络与深层网络的训练与泛化性能。结果捕捉到丰富的现象,包括标度律、双下降及非线性学习动态,为高维深度学习提供了统一的理论视角。
原文摘要 · Abstract (English)
Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down. In particular, the proportional regime where the data dimension, sample size, and number of model parameters are all large and comparable, gives rise to novel and sometimes counterintuitive behaviors. This paper extends traditional Random Matrix Theory (RMT) beyond eigenvalue-based analysis of linear models to address the challenges posed by nonlinear ML models such as DNNs in this regime. We introduce the concept of High-dimensional Equivalent, which unifies and generalizes both Deterministic Equivalent and Linear Equivalent, to systematically address three technical challenges: high dimensionality, nonlinearity, and the need to analyze generic eigenspectral functionals. Leveraging this framework, we provide precise characterizations of the training and generalization performance of linear models, nonlinear shallow networks, and deep networks. Our results capture rich phenomena, including scaling laws, double descent, and nonlinear learning dynamics, offering a unified perspective on the theoretical understanding of deep learning in high dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。