揭示深度注意力网络学习的极限与分层渐进规律
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
- 将深度注意力网络映射为序列多指标模型,建立理论分析框架
- 在高维极限下,精确刻画最优性能与算法性能的临界样本量
- 发现各层学习呈顺序递进,且可在真实场景中观测到
本文研究深度注意力神经网络(由多层共享低秩权重的自注意力层构成)的学习问题。首先将此类模型映射为序列多指标模型——一种推广的多指标模型,适用于序列型协变量。在贝叶斯最优学习设定下,当维度 $D$ 和样本数 $N$ 均趋于无穷且数量级相当的极限条件下,我们推导出最优性能及已知最佳多项式时间算法(近似消息传递)的渐近表现,并确定了实现优于随机预测所需的最小样本复杂度的尖锐阈值。分析表明,各层学习过程具有逐层递进特性。最后,讨论该序列学习现象在实际设置中的可观察性。
原文摘要 · Abstract (English)
In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayesian-optimal learning, in the limit of large dimension $D$ and commensurably large number of samples $N$, we derive a sharp asymptotic characterization of the optimal performance as well as the performance of the best-known polynomial-time algorithm for this setting --namely approximate message-passing--, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。