首个多专家与多头注意力脉冲神经网络3D加速架构,显著提升能效与延迟。
Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers
- 采用3D堆叠集成,实现内存-逻辑与逻辑-逻辑协同设计。
- 相比传统2D CMOS,能效与延迟性能大幅优化。
- 为脑启发式脉冲神经网络提供硬件支持,适合低功耗场景。
脉冲神经网络(SNNs)提供类脑的事件驱动机制,被认为是实现高效能深度学习的关键。混合专家方法模拟神经系统并行分布式处理,引入条件计算策略,在不增加计算操作数的情况下扩展模型容量。此外,脉冲混合专家自注意力机制增强了表征能力,有效捕捉视觉或语言标记间多样化的模式与依赖关系。然而,当前缺乏对脉冲变压器所需高度并行分布式处理的硬件支持。本文提出首个面向混合专家与多头注意力脉冲变压器的3D硬件架构与设计方法。通过采用内存-逻辑与逻辑-逻辑堆叠的3D集成技术,探索具有空间可堆叠电路结构的脑启发式加速器,相比传统2D CMOS集成,显著优化了能效与延迟。
原文摘要 · Abstract (English)
Spiking Neural Networks(SNNs) provide a brain-inspired and event-driven mechanism that is believed to be critical to unlock energy-efficient deep learning. The mixture-of-experts approach mirrors the parallel distributed processing of nervous systems, introducing conditional computation policies and expanding model capacity without scaling up the number of computational operations. Additionally, spiking mixture-of-experts self-attention mechanisms enhance representation capacity, effectively capturing diverse patterns of entities and dependencies between visual or linguistic tokens. However, there is currently a lack of hardware support for highly parallel distributed processing needed by spiking transformers, which embody a brain-inspired computation. This paper introduces the first 3D hardware architecture and design methodology for Mixture-of-Experts and Multi-Head Attention spiking transformers. By leveraging 3D integration with memory-on-logic and logic-on-logic stacking, we explore such brain-inspired accelerators with spatially stackable circuitry, demonstrating significant optimization of energy efficiency and latency compared to conventional 2D CMOS integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。