揭示预测编码网络在无限宽深下的稳定机制,连接其与反向传播的梯度一致性
On the Infinite Width and Depth Limits of Predictive Coding Networks
- 通过非懒惰参数化实现宽深同步缩放下的稳定训练
- 当网络远宽于深时,预测编码梯度趋近反向传播梯度
- 为类脑神经网络中局部更新实现反向传播提供理论依据
预测编码(PC)是一种生物合理替代标准反向传播(BP)的方法,通过优化网络活动来最小化能量函数,再更新权重。近期工作通过引入类BP重参数化提升了深层预测编码网络(PCN)的训练稳定性,但其可扩展性与理论基础仍不明确。本文研究了PCN在无限宽度和深度极限下的行为。对于线性网络,我们推导出在宽深同步缩放下的稳定且非懒惰的参数化形式,发现标准PCN在训练过程中随宽度增加会输出爆炸。在稳定参数化下,我们证明:当网络远宽于深(depth/width → 0)时,活动平衡状态下的PC梯度收敛至BP梯度。实验表明,多种非线性模型(包括卷积网络与Transformer)在大宽度下均表现出高梯度对齐。整体而言,本工作约束了可扩展的预测编码参数化形式,同时暗示在远宽于深的网络(如大脑)中,仅通过局部更新即可实现类似反向传播的计算。
原文摘要 · Abstract (English)
Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of deep PC networks (PCNs) by leveraging some BP-inspired reparameterisations, but the scalability and theoretical basis of these methods remain unclear. To address this gap, we study the infinite width and depth limits of PCNs. For linear networks, we derive stable and "non-lazy" parameterisations when scaling both the model width and depth, revealing that the output of standard PCNs explodes with width during training. Moreover, under stable parameterisations, we show that the gradients computed by PC at activity equilibrium converge to the BP gradients for networks that are much wider than deep ($depth/width\to0$). Experiments show high gradient alignment between PC and BP at large width for different nonlinear models, including convolutional networks and transformers. Overall, this work constrains the parameterisations that are scalable with PC, while suggesting how BP could be implemented using only local updates in much wider than deep networks like the brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。