SpiNNaker2芯片融合脑启发计算与深度学习,实现高效灵活的异构计算。
The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

- 采用152个ARM处理器核与专用加速器,支持事件驱动通信与系统集成。
- 在INT8下达4.5 TOPS算力和2.7 TOPS/W能效,支持超15万神经元仿真。
- 适合研究脑启发计算、低功耗嵌入式系统及深度网络混合架构的开发者。
在深度学习中,效率日益重要以应对模型规模增长。类脑硬件长期被视为替代深度网络的潜在方案,借鉴大脑结构实现前所未有的能效优势。然而,这类优势近期才开始在复杂性和实际应用中显现。SpiNNaker2芯片弥合了深度网络与类脑计算之间的差距,支持灵活探索两者的融合计算方式。该芯片配备152个处理单元,每个单元含ARM M4F处理器和专用加速器,扩展了SpiNNaker路由结构以支持可扩展的事件驱动通信,并提供多种外部接口用于系统集成,包括千兆以太网和LPDDR4内存接口。我们展示了其在类脑计算、深度网络以及新型事件驱动计算任务中的性能与效率。对于深度网络任务,芯片在高性能模式下可达4.5 TOPS算力,在高能效模式下对INT8运算实现2.7 TOPS/W能效。在1毫秒时间步长下,可模拟超过15万神经元和每秒18亿次突触事件。其低于250 mW的基线功耗使其在不同负载条件下仍具能效优势,支持稀疏与事件驱动计算模式的探索。这些结果证明了其作为可扩展类脑计算通用硬件平台的能力,以及与主流深度网络方法结合的潜力。
原文摘要 · Abstract (English)
In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and real-world applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with >150000 neurons and >1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip's capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。