用脉冲神经网络让机器人同时学多任务,省电又高效。
Enabling Energy-Efficient Simultaneous Multi-Task Reinforcement Learning through Spiking Neural Networks with Active Dendrites for Bio-inspired Generalist Agents
- 用带主动树突的脉冲网络动态构建专用子网,应对多任务干扰。
- 在三个Atari游戏上接近人类水平表现,能耗仅为现有方法一半。
- 适合做节能型通用智能体,尤其适用于机器人和无人机场景。
强化学习(RL)在训练自主解决复杂任务的智能体方面表现出色,如移动机器人、无人机/无人车及游戏智能体。然而,同时掌握多个任务(即多任务强化学习)仍面临挑战,这对智能体适应真实环境变化尤为重要。现有方法通过共享神经网络结构提升泛化能力,但存在任务干扰与高能耗问题。为此,我们提出MTSpark,一种基于带主动树突的脉冲神经网络(SNN)的新型多任务强化学习方法,用于构建类生物通用智能体。具体地,MTSpark在深度脉冲Q网络(DSQN)中引入主动树突、双重结构和任务特异性上下文信号,实现对各任务的动态子网构建,并利用稀疏运算提升能效。实验表明,MTSpark在三个Atari游戏上取得优异性能:Pong得分为-5.4,Breakout为0.6,Enduro为371.2,接近人类水平(分别为-3、31、368),且内存相当,能耗约为现有方法的1/2。结果表明,该方法有望推动节能型通用智能体的发展。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated remarkable capabilities in training agents to solve complex tasks autonomously, such as mobile robots, UAVs/UGVs, and game-playing agents). However, scaling RL to master multiple tasks simultaneously (i.e., so-called multi-task RL) remains a significant challenge. Such a multi-task RL capability especially is important for agents to adapt to changes in real-world operational environments. State-of-the-art works show that, training agents with neural networks and shared structures across tasks promises improved generalization in simultaneous multi-task RL. However, they still suffer from task interference and incur high energy consumption due to intensive computation. To address this, we propose MTSpark, a novel methodology that enables energy-efficient simultaneous multi-task RL using spiking neural networks (SNNs) equipped with active dendrites for bio-inspired generalist agents. Specifically, MTSpark enhances a Deep Spiking Q-Network (DSQN) with active dendrites, a dueling structure, and task-specific context signals to dynamically form specialized sub-networks for individual tasks, while exploiting sparse operations for energy-efficient network processing. Experimental results demonstrate that MTSpark achieves higher performance and efficiency compared to state-of-the-art by obtaining high scores across three Atari games (i.e., Pong: -5.4, Breakout: 0.6, and Enduro: 371.2), approaching human-level performance (i.e., Pong: -3, Breakout: 31, Enduro: 368), while incurring similar memory and about 2x lower energy than state-of-the-art. These results show that our MTSpark potentially advances the frontiers toward energy-efficient generalist agents by combining RL and SNNs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。