arXiv:2510.15904cs.ARcs.SY2025-10

用6T SRAM缓存改造出新型存内计算芯片,提升能效并支持深度学习推理。

NVM-in-Cache: Repurposing Commodity 6T SRAM Cache into NVM Analog Processing-in-Memory Engine using a Novel Compute-on-Powerline Scheme

  • 在6T SRAM中嵌入RRAM器件,构建6T-2R混合单元实现存内计算。
  • 在22nm工艺下达到0.4 TOPS算力与452.34 TOPS/W能效,128路并行时对ResNet-18识别准确率达91.76%。
  • 无需增加面积即可扩展存储与计算能力,适合下一代AI加速器和通用处理器。

深度神经网络工作负载的快速增长显著增加了机器学习应用中片上SRAM的大容量需求,如今SRAM阵列已占芯片总面积的相当大比例。为应对存储密度与计算效率双重挑战,本文提出一种将阻变随机存取存储器(RRAM)器件集成到传统6T-SRAM单元中的NVM-in-Cache架构,形成紧凑的6T-2R位单元。该混合单元支持存内计算(PIM)模式,可在缓存供电线上直接执行大规模并行乘加(MAC)操作,同时保留原有缓存数据。利用6T-2R结构的固有特性,该架构实现了额外存储能力,并在无任何位单元面积开销的情况下获得高计算吞吐量。在格罗方德22nm FDSOI工艺下的电路与阵列级仿真表明,所提设计实现0.4 TOPS算力和452.34 TOPS/W能效;对于128行并行操作,通过映射ResNet-18模型完成CIFAR-10分类任务,准确率达到91.76%。这些结果凸显了该方案作为可扩展、高效节能计算方法的潜力,可通过复用现有6T SRAM缓存架构,为下一代AI加速器与通用处理器提供支持。

原文摘要 · Abstract (English)

The rapid growth of deep neural network (DNN) workloads has significantly increased the demand for large-capacity on-chip SRAM in machine learning (ML) applications, with SRAM arrays now occupying a substantial fraction of the total die area. To address the dual challenges of storage density and computation efficiency, this paper proposes an NVM-in-Cache architecture that integrates resistive RAM (RRAM) devices into a conventional 6T-SRAM cell, forming a compact 6T-2R bit-cell. This hybrid cell enables Processing-in-Memory (PIM) mode, which performs massively parallel multiply-and-accumulate (MAC) operations directly on cache power lines while preserving stored cache data. By exploiting the intrinsic properties of the 6T-2R structure, the architecture achieves additional storage capability, high computational throughput without any bit-cell area overhead. Circuit- and array-level simulations in GlobalFoundries 22nm FDSOI technology demonstrate that the proposed design achieves a throughput of 0.4 TOPS and 452.34 TOPS/W. For 128 row-parallel operations, the CIFAR-10 classification is demonstrated by mapping a Resnet-18 neural network, achieving an accuracy of 91.76%. These results highlight the potential of the NVM-in-Cache approach to serve as a scalable, energy-efficient computing method by re-purposing existing 6T SRAM cache architecture for next-generation AI accelerators and general purpose processors.

存内计算非易失内存能效优化神经网络加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。