arXiv:2511.12544cs.ARcs.ET2025-11被引 1

一种可重构的低功耗存内计算宏,支持混合精度推理。

FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration

  • 9T位单元融合存储与异或计算,实现存内运算
  • 22管压缩树加速乘加运算,延迟与功耗更低
  • 4KB宏支持4位浮点/正数格式,能效高达364 TOPS/W

AIoT设备对低功耗、小面积的TinyML推理需求日益增长,亟需减少数据搬移并保持高计算效率的内存架构。本文提出FERMI-ML,一种面向TinyML加速的灵活且资源高效的存内计算(MIS)SRAM宏。该9T XNOR型位单元将5T存储单元与4T XNOR计算单元集成,可在同一阵列中实现可变精度的乘累加(MAC)与内容寻址(CAM)操作。采用22管压缩树结构的累加器,相较传统加法树实现1至64位的对数级乘加计算,显著降低延迟与功耗。4 KB宏支持存内计算与基于CAM的查找功能,兼容Posit-4或FP-4精度。65 nm后版图结果表明,在0.9 V电压下工作频率达350 MHz,吞吐量为1.93 TOPS,能效达364 TOPS/W,InceptionV4与ResNet-18模型的性能保持率(QoR)高于97.5%。因此,FERMI-ML展示了紧凑、可重构、节能的数字存内计算宏,可支持混合精度的TinyML负载。

原文摘要 · Abstract (English)

The growing demand for low-power and area-efficient TinyML inference on AIoT devices necessitates memory architectures that minimise data movement while sustaining high computational efficiency. This paper presents FERMI-ML, a Flexible and Resource-Efficient Memory-In-Situ (MIS) SRAM macro designed for TinyML acceleration. The proposed 9T XNOR-based RX9T bit-cell integrates a 5T storage cell with a 4T XNOR compute unit, enabling variable-precision MAC and CAM operations within the same array. A 22-transistor (C22T) compressor-tree-based accumulator facilitates logarithmic 1-64-bit MAC computation with reduced delay and power compared to conventional adder trees. The 4 KB macro achieves dual functionality for in-situ computation and CAM-based lookup operations, supporting Posit-4 or FP-4 precision. Post-layout results at 65 nm show operation at 350 MHz with 0.9 V, delivering a throughput of 1.93 TOPS and an energy efficiency of 364 TOPS/W, while maintaining a Quality-of-Result (QoR) above 97.5% with InceptionV4 and ResNet-18. FERMI-ML thus demonstrates a compact, reconfigurable, and energy-aware digital Memory-In-Situ macro capable of supporting mixed-precision TinyML workloads.

存内计算TinyML低功耗硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。