揭示循环与深度网络在特征学习中的统一理论
A unified theory of feature learning in RNNs and DNNs
- 基于表示核的平均场理论统一描述RNN与DNN训练过程
- 权重共享使RNN在高信号时产生跨时间步相关表征
- 适用于理解序列任务中RNN的泛化优势
循环神经网络(RNN)和深度神经网络(DNN)是机器学习的核心架构。尽管二者仅通过权重共享区分,可通过时间展开实现等价,但其功能表现却迥异。本文建立了一种统一的平均场理论,从表示核角度刻画完全训练后的RNN与DNN在特征学习(μP)范式下的行为。该理论将训练视为对序列与模式的贝叶斯推断,直接揭示了权重共享带来的功能影响。在典型DNN任务中,我们发现当学习信号超过由随机权重引入的噪声时存在相变:低于阈值时,RNN与DNN行为一致;高于阈值时,仅RNN能生成跨时间步的相关表征。对于序列任务,权重共享还引入归纳偏置,通过插值无监督时间步提升泛化能力。本理论为连接网络结构与功能偏差提供了统一框架。
原文摘要 · Abstract (English)
Recurrent and deep neural networks (RNNs/DNNs) are cornerstone architectures in machine learning. Remarkably, RNNs differ from DNNs only by weight sharing, as can be shown through unrolling in time. How does this structural similarity fit with the distinct functional properties these networks exhibit? To address this question, we here develop a unified mean-field theory for RNNs and DNNs in terms of representational kernels, describing fully trained networks in the feature learning ($μ$P) regime. This theory casts training as Bayesian inference over sequences and patterns, directly revealing the functional implications induced by the RNNs' weight sharing. In DNN-typical tasks, we identify a phase transition when the learning signal overcomes the noise due to randomness in the weights: below this threshold, RNNs and DNNs behave identically; above it, only RNNs develop correlated representations across timesteps. For sequential tasks, the RNNs' weight sharing furthermore induces an inductive bias that aids generalization by interpolating unsupervised time steps. Overall, our theory offers a way to connect architectural structure to functional biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。