揭示深度神经网络学习分层多指标模型的最优缩放规律
Optimal scaling laws in learning hierarchical multi-index models
- 通过信息论推导出子空间恢复与预测误差的精确缩放律
- 发现目标特征按层级顺序通过相变阶段逐步学习,达到最优速率
- 简单谱估计器可实现最优性能,适合研究浅层网络泛化机制
本文针对一类分层多指标目标,为两层神经网络在真正表征受限的环境下提供了精确的缩放律理论。我们推导出子空间恢复和预测误差的严格信息论缩放律,揭示了目标的分层特征如何通过一系列相变阶段逐级学习。进一步表明,这些最优速率可由一种简单、目标无关的谱估计器实现,该估计器可解释为第一层权重梯度下降在小学习率下的极限情形。一旦获得适配表示,读出层即可通过高效过程实现统计最优学习。因此,本工作统一且严谨地解释了浅层神经网络在处理此类分层目标时的缩放律、平台现象与谱结构。
原文摘要 · Abstract (English)
In this work, we provide a sharp theory of scaling laws for two-layer neural networks trained on a class of hierarchical multi-index targets, in a genuinely representation-limited regime. We derive exact information-theoretic scaling laws for subspace recovery and prediction error, revealing how the hierarchical features of the target are sequentially learned through a cascade of phase transitions. We further show that these optimal rates are achieved by a simple, target-agnostic spectral estimator, which can be interpreted as the small learning-rate limit of gradient descent on the first-layer weights. Once an adapted representation is identified, the readout can be learned statistically optimally, using an efficient procedure. As a consequence, we provide a unified and rigorous explanation of scaling laws, plateau phenomena, and spectral structure in shallow neural networks trained on such hierarchical targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。