arXiv:2505.14533cs.LGcs.AI2025-05被引 1

用脉冲神经网络提升强化学习能效,兼顾性能与低功耗。

Energy-Efficient Deep Reinforcement Learning with Spiking Transformers

  • 设计基于多步漏电脉冲神经元的脉冲变压器架构
  • 在多个基准上实现比传统Transformer更高的策略性能
  • 适合资源受限的实时自主系统部署

近年来,基于智能体的Transformer因解决复杂任务的能力而广泛应用。然而,Transformer的高计算复杂性常导致显著能耗,限制其在真实世界自主系统中的部署。脉冲神经网络(SNNs)具有生物启发结构,为机器学习提供了节能替代方案。本文提出一种新型脉冲变压器强化学习(STRL)算法,结合SNN的能效优势与强化学习的强大决策能力。具体而言,设计了一种使用多步漏电积分-放电(LIF)神经元的SNN,并引入注意力机制以处理多时间步的时空模式。该架构进一步通过状态、动作和奖励编码,构建出适用于强化学习任务的类Transformer结构。在前沿基准上的综合数值实验表明,所提出的SNN Transformer相较于传统基于智能体的Transformer实现了显著提升的策略性能。兼具更高能效与策略最优性,本工作为在复杂现实决策场景中部署生物启发的低成本模型指明了新方向。

原文摘要 · Abstract (English)

Agent-based Transformers have been widely adopted in recent reinforcement learning advances due to their demonstrated ability to solve complex tasks. However, the high computational complexity of Transformers often results in significant energy consumption, limiting their deployment in real-world autonomous systems. Spiking neural networks (SNNs), with their biologically inspired structure, offer an energy-efficient alternative for machine learning. In this paper, a novel Spike-Transformer Reinforcement Learning (STRL) algorithm that combines the energy efficiency of SNNs with the powerful decision-making capabilities of reinforcement learning is developed. Specifically, an SNN using multi-step Leaky Integrate-and-Fire (LIF) neurons and attention mechanisms capable of processing spatio-temporal patterns over multiple time steps is designed. The architecture is further enhanced with state, action, and reward encodings to create a Transformer-like structure optimized for reinforcement learning tasks. Comprehensive numerical experiments conducted on state-of-the-art benchmarks demonstrate that the proposed SNN Transformer achieves significantly improved policy performance compared to conventional agent-based Transformers. With both enhanced energy efficiency and policy optimality, this work highlights a promising direction for deploying bio-inspired, low-cost machine learning models in complex real-world decision-making scenarios.

强化学习脉冲神经网络能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。