arXiv:2410.13991math.STcs.LG2024-10被引 2

解析带尖峰协方差的线性模型泛化误差,揭示主成分影响机制。

Generalization for Least Squares Regression With Simple Spiked Covariances

  • 基于随机矩阵理论分析隐藏层特征矩阵谱结构
  • 在渐近比例下推导出尖峰协方差模型的泛化误差
  • 发现尖峰对应的特征向量与特征值决定泛化性能

随机矩阵理论已被证明是分析线性模型泛化能力的有力工具。然而,即使是对梯度下降训练的两层神经网络,其泛化性质仍不清晰。要理解此类网络的泛化表现,关键在于刻画隐藏层特征矩阵的谱结构。近期研究通过描述单步梯度更新后的谱结构,揭示了尖峰协方差特性。然而,带有尖峰协方差的线性模型的泛化误差此前尚未被确定。本文针对两类具有尖峰协方差的简单模型,研究其在渐近比例情形下的泛化误差。分析表明,尖峰对应的特征值与特征向量对泛化误差有显著影响。

原文摘要 · Abstract (English)

Random matrix theory has proven to be a valuable tool in analyzing the generalization of linear models. However, the generalization properties of even two-layer neural networks trained by gradient descent remain poorly understood. To understand the generalization performance of such networks, it is crucial to characterize the spectrum of the feature matrix at the hidden layer. Recent work has made progress in this direction by describing the spectrum after a single gradient step, revealing a spiked covariance structure. Yet, the generalization error for linear models with spiked covariances has not been previously determined. This paper addresses this gap by examining two simple models exhibiting spiked covariances. We derive their generalization error in the asymptotic proportional regime. Our analysis demonstrates that the eigenvector and eigenvalue corresponding to the spike significantly influence the generalization error.

泛化误差随机矩阵线性模型尖峰协方差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。