提出抗错自旋内存架构,实现高效低延迟的神经网络计算。
CRAM-ER: Error-Resilient Spintronic Computational Random Access Memory for Scalable In-Memory Computation

- 采用自旋电子学与CMOS混合加法树结构,降低器件误差影响。
- 在多个DNN基准上实现近无损精度,延迟降低两个数量级。
- 适合需要高能效、低延迟的神经网络加速场景。
深度神经网络(DNN)在多个领域达到顶尖性能,但传统冯·诺依曼计算范式面临严重的内存瓶颈。新兴的近内存和存内计算方法虽缓解此问题,但引入显著外围开销。基于磁性随机存取存储器(MRAM)的计算随机存取内存(CRAM)可在无外围开销下实现原位逻辑运算,提供高密度、低功耗解决方案。然而,概率性MRAM开关导致门级错误,限制了CRAM在加速DNN时的可扩展性和可靠性;且大量串行MRAM写入严重制约了CRAM吞吐量。为此,我们提出一种面向可扩展存内矩阵-向量乘法(MVM)的抗错CRAM(CRAM-ER)架构。通过误差感知软硬件协同设计框架,利用混合自旋电子学-CRAM + CMOS加法树架构减轻器件级误差影响,实现了高面积与能量效率的MVM功能。进一步开发误差感知模型微调与细粒度误差校正技术以增强抗错能力。在DNN基准上的评估表明,该混合架构实现近无损精度,同时将CRAM延迟降低达两个数量级,在能效与能量延迟积方面优于CPU/GPU+高带宽DRAM。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) have achieved state-of-the-art performance across diverse domains. However, typical Von Neumann compute paradigms face severe memory bottlenecks. Emerging near-memory and compute-in-memory approaches alleviate this but incur significant peripheral overhead. Computational Random Access Memory (CRAM) based on MRAM enables in-situ logic without peripheral overhead, offering a dense, energy-efficient solution. However, probabilistic MRAM switching induces gate-level errors that limit the scalability and reliability of CRAM for accelerating DNN. Moreover, the large number of sequential MRAM writes severely constrains CRAM throughput. To address these challenges, we propose an error-resilient CRAM (CRAM-ER) architecture for scalable in-memory matrix-vector multiplications (MVMs). Our error-aware hardware-software co-design framework leverages a hybrid spintronic-CRAM + CMOS adder-tree architecture to mitigate the impact of device-level errors, demonstrating MVM functionality with high area and energy efficiency. We further develop an error-aware model fine-tuning and fine-grained error correction for enhanced error resilience. Evaluations of the CMOS+spintronic hybrid architecture on DNN benchmarks show near-lossless accuracy while reducing CRAM latency by up to 2 orders of magnitude, outperforming CPU/GPU+high-bandwidth DRAM in both energy efficiency and energy-delay product.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。