arXiv:2511.02401math.STcs.LG2025-11

用随机矩阵理论分析循环网络泛化能力,揭示其对近期输入的偏好。

Generalization in Representation Models via Random Matrix Theory: Application to Recurrent Networks

  • 基于随机矩阵理论推导出固定特征模型的泛化误差公式。
  • 发现线性ESN等价于带指数加权输入协方差的岭回归,具短期记忆偏好。
  • 实验验证:小样本短记忆时ESN更优,数据多或长依赖时岭回归胜出。

我们研究了使用固定特征表示(冻结中间层)后接可训练读出层的模型的泛化误差。该设置涵盖从深度随机特征模型到具有循环动态的回声状态网络(ESNs)等多种架构。在高维情形下,运用随机矩阵理论推导出渐近泛化误差的闭式表达式,并将其应用于循环表示,得到简洁性能刻画公式。令人惊讶的是,我们证明线性ESN等价于带指数时间加权(‘记忆’)输入协方差的岭回归,揭示出对近期输入的明确归纳偏置。实验结果与预测一致:在小样本、短记忆场景中ESN表现更优,而在数据充足或存在长程依赖时岭回归更具优势。该方法为分析过参数化模型提供了通用框架,并深化了对深度学习网络行为的理解。

原文摘要 · Abstract (English)

We first study the generalization error of models that use a fixed feature representation (frozen intermediate layers) followed by a trainable readout layer. This setting encompasses a range of architectures, from deep random-feature models to echo-state networks (ESNs) with recurrent dynamics. Working in the high-dimensional regime, we apply Random Matrix Theory to derive a closed-form expression for the asymptotic generalization error. We then apply this analysis to recurrent representations and obtain concise formula that characterize their performance. Surprisingly, we show that a linear ESN is equivalent to ridge regression with an exponentially time-weighted (''memory'') input covariance, revealing a clear inductive bias toward recent inputs. Experiments match predictions: ESNs win in low-sample, short-memory regimes, while ridge prevails with more data or long-range dependencies. Our methodology provides a general framework for analyzing overparameterized models and offers insights into the behavior of deep learning networks.

循环网络泛化分析随机矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。