将预测编码与最小描述长度结合,为深度学习提供理论支撑
Bridging Predictive Coding and MDL: A Two-Part Code Framework for Deep Learning
- 用分块坐标下降法连接预测编码与最小描述长度目标
- 推导出含压缩项的泛化误差上界,证明训练过程持续优化
- 首次给出预测编码的收敛性与泛化保证,适合关注理论机制的研究者
本文首次建立预测编码(PC)与最小描述长度(MDL)原则在深度网络中的理论联系。证明逐层PC执行的是MDL二部分码目标的分块坐标下降,同时最小化经验风险与模型复杂度。利用霍夫丁不等式和前缀码先验,推导出形式为 $R(θ) \ le \ hat{R}(θ) + \ frac{L(θ)}{N}$ 的新泛化界,刻画了拟合与压缩之间的权衡。进一步证明每次PC扫描单调减少经验二部分码长,得到比无约束梯度下降更紧的高概率风险界。最后,表明重复PC更新收敛至分块坐标驻点,逼近MDL最优解。这是首个为PC训练的深度模型提供正式泛化与收敛保证的结果,使预测编码成为具有理论基础且符合生物可解释性的反向传播替代方案。
原文摘要 · Abstract (English)
We present the first theoretical framework that connects predictive coding (PC), a biologically inspired local learning rule, with the minimum description length (MDL) principle in deep networks. We prove that layerwise PC performs block-coordinate descent on the MDL two-part code objective, thereby jointly minimizing empirical risk and model complexity. Using Hoeffding's inequality and a prefix-code prior, we derive a novel generalization bound of the form $R(θ) \le \hat{R}(θ) + \frac{L(θ)}{N}$, capturing the tradeoff between fit and compression. We further prove that each PC sweep monotonically decreases the empirical two-part codelength, yielding tighter high-probability risk bounds than unconstrained gradient descent. Finally, we show that repeated PC updates converge to a block-coordinate stationary point, providing an approximate MDL-optimal solution. To our knowledge, this is the first result offering formal generalization and convergence guarantees for PC-trained deep models, positioning PC as a theoretically grounded and biologically plausible alternative to backpropagation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。