证明了脉冲网络能通过梯度学习时间编码信息,且延迟可提升性能。
Beyond Rate Coding: Surrogate Gradients Enable Spike Timing Learning in Spiking Neural Networks
- 用替代梯度法训练脉冲网络,挖掘脉冲时间而非发放率的编码信息。
- 在移除发放率信息的语音数据上仍保持高精度,证明其利用了时间编码。
- 带可训练轴突延迟的网络更擅长捕捉时空结构,适合时序任务研究者。
替代梯度下降算法使脉冲神经网络能够完成复杂的感知任务,是理解脉冲如何参与神经计算的重要一步。然而,该算法是否充分探索了脉冲解码的可能性尚不明确。我们探究了基于替代梯度训练的脉冲网络能否利用仅编码于脉冲时间而非发放率的信息。构建了包含脉冲间隔、时空脉冲模式(多时程性)和同步编码等类型的时间信息的合成数据集,发现替代梯度训练能有效提取所有这些信息。在更真实的语音数据集上,同时存在时间和发放率信息。为此,我们构造了移除全部发放率信息的变体数据集,发现替代梯度训练仍表现良好。我们在有无可训练轴突延迟的情况下测试了网络性能,发现延迟显著提升性能,尤其在挑战性任务中。为探究语音任务中网络使用的时间编码类型,我们对脉冲进行时间倒置,破坏时空模式但保留脉冲间隔和同步信息。结果表明:无延迟网络在时间倒置下表现稳定,而含延迟网络性能显著下降,说明带延迟的脉冲网络更能利用时间结构。为促进后续研究,我们已发布修改后的语音数据集。
原文摘要 · Abstract (English)
The surrogate gradient descent algorithm enabled spiking neural networks to be trained to carry out challenging sensory processing tasks, an important step in understanding how spikes contribute to neural computations. However, it is unclear the extent to which these algorithms fully explore the space of possible spiking solutions to problems. We investigated whether spiking networks trained with surrogate gradient descent can learn to make use of information that is only encoded in the timing and not the rate of spikes. We constructed synthetic datasets with a range of types of spike timing information (interspike intervals, spatio-temporal spike patterns or polychrony, and coincidence codes). We find that surrogate gradient descent training can extract all of these types of information. In more realistic speech-based datasets, both timing and rate information is present. We therefore constructed variants of these datasets in which all rate information is removed, and find that surrogate gradient descent can still perform well. We tested all networks both with and without trainable axonal delays. We find that delays can give a significant increase in performance, particularly for more challenging tasks. To determine what types of spike timing information are being used by the networks trained on the speech-based tasks, we test these networks on time-reversed spikes which perturb spatio-temporal spike patterns but leave interspike intervals and coincidence information unchanged. We find that when axonal delays are not used, networks perform well under time reversal, whereas networks trained with delays perform poorly. This suggests that spiking neural networks with delays are better able to exploit temporal structure. To facilitate further studies of temporal coding, we have released our modified speech-based datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。