arXiv:2510.24461cs.AIcs.RO2025-10NeurIPS被引 1

改进脉冲神经网络的梯度方法,让机器人控制更高效稳定

Adaptive Surrogate Gradients for Sequential Reinforcement Learning in Spiking Neural Networks

  • 设计可自适应调整的代理梯度斜率,提升深层网络梯度强度
  • 结合引导策略训练,使无人机控制平均得分达400分,超前人方法2倍
  • 适用于需低功耗、实时处理的神经形态机器人系统

类脑计算系统有望通过数个数量级的能效提升,推动能源受限机器人发展,并实现原生时间处理。脉冲神经网络(SNN)是此类系统的有前途算法,但其在复杂控制任务中的应用面临两大挑战:(1) 脉冲神经元不可微,需依赖代理梯度,但其优化性质不明确;(2) SNN具有状态动态特性,需序列训练,而强化学习中早期训练序列过短,难以跨越启动期。本文系统分析代理梯度斜率设置,发现较浅斜率虽增强深层梯度幅度,但降低与真实梯度的对齐性。在监督学习中,固定或调度斜率无明显优劣;但在强化学习中,较浅或调度斜率使训练和最终部署性能均提升2.1倍。进一步提出新训练方法,利用特权引导策略启动学习,同时保留在线环境交互。结合自适应斜率调度,在真实无人机位置控制任务中实现平均回报400分,显著优于行为克隆和TD3BC等先前方法(最高仅-200分)。本工作推进了对SNN代理梯度学习的理论理解,并验证了神经形态控制器在真实机器人系统中的实用性。

原文摘要 · Abstract (English)

Neuromorphic computing systems are set to revolutionize energy-constrained robotics by achieving orders-of-magnitude efficiency gains, while enabling native temporal processing. Spiking Neural Networks (SNNs) represent a promising algorithmic approach for these systems, yet their application to complex control tasks faces two critical challenges: (1) the non-differentiable nature of spiking neurons necessitates surrogate gradients with unclear optimization properties, and (2) the stateful dynamics of SNNs require training on sequences, which in reinforcement learning (RL) is hindered by limited sequence lengths during early training, preventing the network from bridging its warm-up period. We address these challenges by systematically analyzing surrogate gradient slope settings, showing that shallower slopes increase gradient magnitude in deeper layers but reduce alignment with true gradients. In supervised learning, we find no clear preference for fixed or scheduled slopes. The effect is much more pronounced in RL settings, where shallower slopes or scheduled slopes lead to a 2.1x improvement in both training and final deployed performance. Next, we propose a novel training approach that leverages a privileged guiding policy to bootstrap the learning process, while still exploiting online environment interactions with the spiking policy. Combining our method with an adaptive slope schedule for a real-world drone position control task, we achieve an average return of 400 points, substantially outperforming prior techniques, including Behavioral Cloning and TD3BC, which achieve at most --200 points under the same conditions. This work advances both the theoretical understanding of surrogate gradient learning in SNNs and practical training methodologies for neuromorphic controllers demonstrated in real-world robotic systems.

脉冲神经网络强化学习类脑计算机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。