用新型三值脉冲神经元提升深度强化学习性能
Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons
- 设计三值脉冲神经元缓解梯度估计偏差问题
- 在7个Atari游戏上平均得分超越二值神经元基线
- 适合研究脉冲神经网络与强化学习融合的学者
我们提出一种新的三值脉冲神经元模型,以提升二值脉冲神经元在深度Q学习中的表征能力。尽管近期已有三值神经元模型被提出以克服二值神经元的表征局限,但我们发现其在深度Q学习任务中的表现反而劣于二值模型,与先前研究结论相悖。通过数学与实证分析,我们推测训练过程中的梯度估计偏差是根本原因。所提出的三值脉冲神经元通过降低估计偏差来缓解该问题。我们将该神经元作为核心计算单元构建深度脉冲Q学习网络,命名为深度非对称三值脉冲Q网络(DATSQN),并在Gym环境的7个Atari游戏中评估其性能。结果表明,该模型有效缓解了三值神经元在DQN任务中的性能下降,并在本文设定的评估条件下,相较于二值基线提升了平均游戏得分。
原文摘要 · Abstract (English)
We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has recently been introduced to overcome the limited representation capacity offered by binary spiking neurons, we show that its performance is worse than that of binary models in deep Q-learning tasks, contradicting previous findings from recent studies. Through mathematical and empirical analysis, we hypothesize that gradient estimation bias during training is the underlying cause. The proposed ternary spiking neuron model mitigates this issue by reducing the estimation bias. We use the proposed ternary spiking neuron as the fundamental computing unit in a deep spiking Q-learning network, which we call the deep asymmetric ternary spiking Q-network (DATSQN), and evaluate the network's performance in seven Atari games from the Gym environment. The results show that the proposed ternary spiking neuron mitigates the performance degradation of ternary neurons in DQN tasks and improves the mean game score relative to the binary baseline under the evaluation settings used in this paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。