用赛道内存实现嵌入式CNN加速,提升能效与密度
Hardware-software co-exploration with racetrack memory based in-memory computing for CNN inference in embedded systems
- 设计专用内存计算单元,支持卷积运算的乘累加操作
- 在面积和能耗约束下,实现能效与性能显著提升
- 适合资源受限的嵌入式AI系统,尤其关注低功耗场景
深度神经网络产生并处理大量数据,给低资源嵌入式系统带来挑战。存内计算已被证明是一种高效的计算架构,适用于嵌入式AI应用。在新兴存储技术中,赛道内存是一种非易失性技术,具备高数据密度特性,适合存内计算。然而,将存内算术电路与存储单元集成会影响存储密度和能效。如何在面积和能耗限制下构建高效的赛道内存存内算术电路仍具挑战。为此,我们提出一种面向赛道内存优化的高效存内卷积神经网络(CNN)加速器。设计了一系列适配乘累加操作的存内计算单元。同时,探索基于赛道内存的系统与CNN模型架构的设计空间,采用软硬件协同设计,在保持模型精度的同时,提升赛道内存嵌入式系统的效率与性能。所设计电路与模型-系统协同优化策略在小内存阵列面积下,显著提升了能效与性能。
原文摘要 · Abstract (English)
Deep neural networks generate and process large volumes of data, posing challenges for low-resource embedded systems. In-memory computing has been demonstrated as an efficient computing infrastructure and shows promise for embedded AI applications. Among newly-researched memory technologies, racetrack memory is a non-volatile technology that allows high data density fabrication, making it a good fit for in-memory computing. However, integrating in-memory arithmetic circuits with memory cells affects both the memory density and power efficiency. It remains challenging to build efficient in-memory arithmetic circuits on racetrack memory within area and energy constraints. To this end, we present an efficient in-memory convolutional neural network (CNN) accelerator optimized for use with racetrack memory. We design a series of fundamental arithmetic circuits as in-memory computing cells suited for multiply-and-accumulate operations. Moreover, we explore the design space of racetrack memory based systems and CNN model architectures, employing co-design to improve the efficiency and performance of performing CNN inference in racetrack memory while maintaining model accuracy. Our designed circuits and model-system co-optimization strategies achieve a small memory bank area with significant improvements in energy and performance for racetrack memory based embedded systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。