用脉冲神经网络实现连续控制强化学习,性能接近传统模型。
Spiking Neural Networks for Continuous Control: Neuromorphic Reinforcement Learning in Conventional Computing

- 设计SANSAC框架,将脉冲神经网络用于连续控制任务
- 在常规计算机上验证其性能与SAC几乎相当
- 为未来类脑硬件上的强化学习研究提供基准
过去十年中,强化学习(RL)算法已在多种问题和控制任务中取得进展。然而,在类脑硬件上部署用于连续控制的RL仍缺乏充分验证。具体而言,尚不清楚将传统的策略网络替换为脉冲神经网络(SNN)是否会影响智能体性能,尤其是在硬件特有优势显现前。本文系统验证了基于软演员-评论家(SAC)的最小化、类脑可行的脉冲策略变体在传统硬件上的表现,建立了未来类脑强化学习研究的基准。我们提出了脉冲策略网络软演员-评论家(SANSAC),专为连续环境中的强化学习框架设计,可适配类脑硬件。在传统计算机上,我们将传统SAC与SANSAC进行对比,结果表明二者性能接近,同时分析了隐藏层维度的影响。实验验证了基于SNN的算法在复杂连续环境中的可行性,且在传统计算平台上表现出与传统神经网络相当的竞争力,为继续探索SNN在连续强化学习框架中的应用提供了基础。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms have made strides over the past decade applying them to a wide range of problems and control tasks. However, the deployment of RL on neuromorphic hardware for continuous control tasks remains under-validated. Namely it is unclear whether replacing a conventional actor network with a spiking neural network (SNN) affects the performance of an agent before any hardware-specific benefits manifest. We provide a systematic validation of a minimal, neuromorphically viable spiking actor variant of Soft Actor-Critic (SAC) on conventional hardware, establishing a baseline for future neuromorphic RL research. In this paper, we propose the Spiking Actor Network Soft Actor Critic (SANSAC) to address the use of RL frameworks in continuous environments, designed as a framework that can be implemented on neuromorphic hardware. We compare a traditional Soft Actor Critic (SAC) network to SANSAC in a traditional computer. We demonstrate the near equivalent performance of SANSAC and SAC, while addressing the impact of hidden dimensions. Our results demonstrate the viability of SNN based algorithms in complex continuous environments, as well as competitive performance to traditional neural networks in traditional computers, providing a basis to continue exploring the use of SNNs in continuous RL frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。