arXiv:2410.18858cond-mat.dis-nncs.LG2024-10被引 3

提出双线性序列回归模型,解析长序列高维标记学习的理论极限。

Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens

  • 构建双线性序列回归模型,模拟高维标记序列的学习机制。
  • 计算出贝叶斯最优泛化误差,并设计匹配性能的消息传递算法。
  • 揭示梯度下降在该模型中的反直觉行为,适合理论研究者阅读。

当前人工智能进展主要围绕大规模语言模型展开,这些模型处理由高维向量(称为标记)组成的长序列。统计物理为神经网络学习机制提供了强大工具,并在现代机器学习发展中发挥了重要作用。然而,针对长序列高维标记的简化且可解析的模型仍研究不足。受单层教师-学生感知机(即广义线性回归)在全连接网络理论中关键作用的启发,本文引入并研究了双线性序列回归(BSR)模型,作为令牌序列最基本的模型之一。我们注意到,现代架构因跳跃连接自然包含了BSR模型。基于近期方法论进展,我们在长序列高维标记极限下计算了该模型的贝叶斯最优泛化误差,并提出了达到此性能的消息传递算法。我们量化了最优学习相较于将序列向量化后采用简单线性回归所带来的改进。此外,还揭示了梯度下降算法在该模型中的令人惊讶的特性。

原文摘要 · Abstract (English)

Current progress in artificial intelligence is centered around so-called large language models that consist of neural networks processing long sequences of high-dimensional vectors called tokens. Statistical physics provides powerful tools to study the functioning of learning with neural networks and has played a recognized role in the development of modern machine learning. The statistical physics approach relies on simplified and analytically tractable models of data. However, simple tractable models for long sequences of high-dimensional tokens are largely underexplored. Inspired by the crucial role models such as the single-layer teacher-student perceptron (aka generalized linear regression) played in the theory of fully connected neural networks, in this paper, we introduce and study the bilinear sequence regression (BSR) as one of the most basic models for sequences of tokens. We note that modern architectures naturally subsume the BSR model due to the skip connections. Building on recent methodological progress, we compute the Bayes-optimal generalization error for the model in the limit of long sequences of high-dimensional tokens, and provide a message-passing algorithm that matches this performance. We quantify the improvement that optimal learning brings with respect to vectorizing the sequence of tokens and learning via simple linear regression. We also unveil surprising properties of the gradient descent algorithms in the BSR model.

序列建模统计物理理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。