提出新型硬件结构,让神经网络延迟更省内存、更高效。
Efficient Synaptic Delay Implementation in Digital Event-Driven AI Accelerators
- 用共享环形延迟队列实现突触延迟,降低内存占用。
- 内存开销随模型稀疏度变化,而非单纯依赖网络规模。
- 适合做软硬件协同优化,尤其在低功耗场景下表现优异。
神经网络中的突触延迟参数化虽未被充分探索,但近期研究表明,这类模型通常更简单、更小、更稀疏,因此在同等任务精度下比非延迟模型更节能。本文提出一种新型硬件结构——共享环形延迟队列(SCDQ),用于支持数字脉冲类脑加速器中的突触延迟。分析与硬件实测表明,该结构在内存占用方面优于现有主流方法,且内存扩展可由模型稀疏度调节,而非仅取决于网络大小。此外,还评估了其在延迟、面积和单次推理能耗方面的性能表现。
原文摘要 · Abstract (English)
Synaptic delay parameterization of neural network models have remained largely unexplored but recent literature has been showing promising results, suggesting the delay parameterized models are simpler, smaller, sparser, and thus more energy efficient than similar performing (e.g. task accuracy) non-delay parameterized ones. We introduce Shared Circular Delay Queue (SCDQ), a novel hardware structure for supporting synaptic delays on digital neuromorphic accelerators. Our analysis and hardware results show that it scales better in terms of memory, than current commonly used approaches, and is more amortizable to algorithm-hardware co-optimizations, where in fact, memory scaling is modulated by model sparsity and not merely network size. Next to memory we also report performance on latency area and energy per inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。