揭示大型线性自编码器的五种极端学习模式及其几何结构。
A prism hierarchy of learning regimes in large linear autoencoders
- 通过损失展开层次,将学习模式与三角双锥几何面对应。
- 在四种模式下推导出训练与泛化损失的显式演化公式。
- 适用于研究深度学习理论极限的科研人员。
机器学习理论研究常关注梯度下降在不同极限情形下的可分析性。然而,系统地刻画特定模型所有定性不同的极端学习模式仍具挑战。本文针对由输入维数、潜在维数、初始化幅度和训练集大小定义的大规模权重共享线性自编码器,提出此类模式的系统图像。该模型在权重上为非线性,其梯度流无通用解析解。我们发现,在形式损失展开层次中,其极端模式自然对应于一个三角双锥的面。具体有五个基本极端模式,分别对应双面:(1) 大数据、(2) 小数据、(3) 均值场、(4) 狭窄潜在空间、(5) 自由。对前四种模式,我们推导出梯度流下训练与总体损失演化的显式表达式,与实验结果高度一致。
原文摘要 · Abstract (English)
Theoretical studies of machine learning models commonly consider different limiting regimes in which the learning dynamics of gradient descent becomes theoretically tractable. It is, however, desirable to have a systematically obtained picture of all qualitatively different extreme learning regimes for a particular type of models. In this paper we propose such a picture for large weight-tied linear autoencoders characterized by input and latent dimensions, initialization magnitude, and training set size. This model is nonlinear in the weights and its gradient flow does not have a general theoretical solution. We show that at the level of the formal loss-expansion hierarchy, its extreme regimes are naturally associated with faces of a triangular prism. In particular, there are five basic extreme regimes associated with the 2-faces of the prism: (1) large-data, (2) small-data, (3) mean-field, (4) narrow-latent, and (5) free. For regimes (1,2,3,4), we derive explicit expressions for both train and population limiting loss evolutions under gradient flow, obtaining very good agreement with experimental results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。