arXiv:2605.20717cs.NEcs.AR2026-05

超低功耗存内计算芯片,支持传统与脉冲神经网络推理。

E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

论文配图:E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference
图 1 · 摘自论文原文
  • 采用3T1R阻变存储器设计,实现高效存内乘法运算。
  • 能效达419 TOPS/W,延迟低至0.48纳秒,精度保持97%以上。
  • 适合边缘AI、物联网及生物传感等低功耗场景。

本文提出E-ReCON,一种基于紧凑3T1R ReRAM单元的16 Kb数字存内计算(DCIM)宏,适用于边缘AI推理。该单元面积仅0.85 μm²,可可靠实现基于AND的存内乘法,兼容传统卷积神经网络(CNN)与脉冲神经网络(SNN)任务。为降低累加开销,引入新型交错式10T/28T加法树,相比传统28T RCA设计,晶体管数量减少37%,功耗降低28%。在65 nm CMOS工艺下,工作电压1.2 V时,最低延迟0.48 ns,吞吐量2.31–3.1 TOPS,能效最高达419 TOPS/W。在LeNet-5、AlexNet和CNN-8模型上,于MNIST/A-Z、CIFAR10和SVHN数据集上分别取得97.81%、93.23%和96.51%的准确率。40%剪枝后,保留近99.8%原始精度,显著降低MAC操作与计算周期。对于SNN任务,该AND型单元支持低开关活动的脉冲-权重乘法,2A2W配置在VGG-8、VGG-16和ResNet-18网络上于CIFAR-10、CIFAR-100和ImageNet-1K数据集上逼近FP32基线精度。相比先前基于ADC的ReRAM-CIM设计,本架构在全PVT与ReRAM变异条件下仍保持鲁棒性,延迟与能效提升近30%-40%。整体上,E-ReCON为下一代边缘AI、物联网、生物传感与类脑计算提供可扩展、低延迟、高能效平台。

原文摘要 · Abstract (English)

This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference. The proposed bitcell occupies only 0.85 um^2 and supports reliable AND-based in-memory multiplication for both conventional convolutional neural network (CNN) and spiking neural network (SNN) workloads. To reduce accumulation overhead, a novel interleaved 10T/28T adder tree is introduced, reducing transistor count and power consumption by 37% and 28%, respectively, compared to a conventional 28T RCA-based design. Implemented in 65 nm CMOS at 1.2 V, the proposed macro achieves a minimum latency of 0.48 ns, throughput of 2.31-3.1 TOPS, and energy efficiency of up to 419 TOPS/W. When evaluated on LeNet-5, AlexNet, and CNN-8 models, the macro achieves 97.81%, 93.23%, and 96.51% accuracy on MNIST/A-Z, CIFAR10, and SVHN datasets, respectively. In addition, 40% pruning preserves nearly 99.8% of the original accuracy while reducing MAC operations and computation cycles. For SNN-oriented workloads, the proposed AND-type bitcell efficiently supports spike-weight multiplication with low switching activity, where the 2A2W configuration achieves accuracy close to the FP32 baseline across VGG-8, VGG-16, and ResNet-18 networks on CIFAR-10, CIFAR-100, and ImageNet-1K datasets. Compared to prior ADC-based ReRAM-CIM designs, the proposed architecture improves latency and energy efficiency by nearly 30-40% while maintaining robust operation under full PVT and ReRAM variability. Overall, E-ReCON provides a scalable, low-latency, and energy-efficient nvCIM platform for next-generation edge-AI, IoT, biomedical sensing, and neuromorphic applications.

存内计算边缘智能阻变存储器神经网络硬件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。