提出新型内存计算架构,用随机乘法提升矩阵运算能效
OISMA: On-the-fly In-memory Stochastic Multiplication Architecture for Matrix-Multiplication Workloads
- 将内存读取转为原位随机乘法,利用准随机计算简化运算
- 180nm实现0.789 TOPS/W能效,22nm可提升百倍能效
- 适合大规模矩阵运算场景,尤其对算力与能效敏感的AI应用
当前人工智能模型复杂度持续上升,大规模矩阵乘法成为主要计算瓶颈。内存计算(IMC)架构被提出以突破冯·诺依曼瓶颈,但数字/二值和模拟IMC架构均存在性能与能效限制。本文提出OISMA,一种基于准随机计算(SC)域的高效内存计算架构,采用弯金字塔(BP)系统,在保持数字存储效率、可扩展性与生产力的同时,实现计算简化。OISMA将常规内存读取操作转换为原位随机乘法,通过外围累加单元处理输出比特流,完成矩阵乘法功能。采用商用180nm工艺与自研阻变存储器(RRAM)构建4-kB 1T1R OISMA阵列。在50 MHz下,能效达0.789 TOPS/W,面积效率3.98 GOPS/mm²,有效计算面积0.804241 mm²。在22nm节点下,相比密集矩阵乘法IMC架构,能效提升两个数量级,面积效率提升一个数量级。
原文摘要 · Abstract (English)
Artificial intelligence (AI) models are currently driven by a significant upscaling of their complexity, with massive matrix-multiplication workloads representing the major computational bottleneck. In-memory computing (IMC) architectures are proposed to avoid the von Neumann bottleneck. However, both digital/binary-based and analog IMC architectures suffer from various limitations, which significantly degrade the performance and energy efficiency gains. This work proposes OISMA, an energy-efficient IMC architecture that utilizes the computational simplicity of a quasi-stochastic computing (SC) domain (bent-pyramid (BP) system) while keeping the same efficiency, scalability, and productivity of digital memories. OISMA converts normal memory read operations into in situ stochastic multiplication operations with a negligible cost. An accumulation periphery then accumulates the output multiplication bitstreams, achieving the matrix multiplication (MatMul) functionality. A 4-kB 1T1R OISMA array was implemented using a commercial 180-nm technology node and in-house resistive random-access memory (RRAM) technology. At 50 MHz, it achieves 0.789 TOPS/W and 3.98 GOPS/mm2 for energy and area efficiency, respectively, occupying an effective computing area of 0.804241 mm2. Scaling OISMA to 22-nm technology shows a significant improvement of two orders of magnitude in energy efficiency and one order of magnitude in area efficiency, compared to dense MatMul IMC architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。