arXiv:2508.15989cs.LGcs.ET2025-08被引 1

通过中间误差信号提升深度卷积CRNN的平衡传播训练能力

Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs

  • 引入分层学习信号作为辅助监督,缓解深层网络梯度消失
  • 在CIFAR-10/100上实现VGG级深度模型的先进性能
  • 首次结合知识蒸馏与局部误差信号,推动EP向深层应用扩展

平衡传播(EP)是一种受生物学启发的局部学习规则,最初用于收敛型循环神经网络(CRNN),其突触更新仅依赖于两个不同阶段的神经元状态。EP计算的梯度与反向传播时序(BPTT)高度一致,同时显著降低计算开销,使其成为类脑架构片上训练的潜在候选方案。然而,以往研究受限于浅层结构,深层网络因梯度消失问题导致能量最小化与梯度计算难以收敛。为此,本文提出一种新型EP框架,引入逐层学习信号作为辅助监督,增强神经元动力学的收敛性。这是首个将知识蒸馏与局部误差信号融入EP的工作,实现了对更深层次架构的训练。所提方法在CIFAR-10和CIFAR-100数据集上达到当前最优表现,验证了其在深层VGG架构上的可扩展性。该成果显著推进了EP的可扩展性,表明中间学习信号可有效拓展EP在深层网络中的实际应用。

原文摘要 · Abstract (English)

Equilibrium Propagation (EP) is a biologically inspired local learning rule first proposed for convergent recurrent neural networks (CRNNs), in which synaptic updates depend only on neuron states from two distinct phases. EP estimates gradients that closely align with those computed by Backpropagation Through Time (BPTT) while significantly reducing computational demands, positioning it as a potential candidate for on-chip training in neuromorphic architectures. However, prior studies on EP have been constrained to shallow architectures, as deeper networks suffer from the vanishing gradient problem, leading to convergence difficulties in both energy minimization and gradient computation. To alleviate the vanishing gradient problem in deep EP networks, we propose a novel EP framework that incorporates layer-wise learning signals to provide auxiliary supervision, which enhances the convergence of neuron dynamics. This is the first work to integrate knowledge distillation and local error signals into EP, enabling the training of significantly deeper architectures. Our proposed approach achieves state-of-the-art performance on the CIFAR-10 and CIFAR-100 datasets, showcasing its scalability on deep VGG architectures. These results represent a significant advancement in the scalability of EP, suggesting that intermediate learning signals can extend the practical applicability of EP to deeper architectures.

平衡传播深度学习类脑计算梯度消失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。