arXiv:2510.19260cs.ARcs.ET2025-10

提出可共享资源的数字存算一体单元,提升边缘AI芯片能效与扩展性。

Res-DPU: Resource-shared Digital Processing-in-memory Unit for Edge-AI Workloads

  • 采用双端口5T SRAM和共享2T AND门,降低每比特乘法晶体管数至5.25T
  • 16KB宏单元达0.43 TOPS吞吐、87.22 TOPS/W能效,ResNet-18在CIFAR-10上达96.85%精度
  • 支持运行时精度-延迟权衡,无需纠错电路,适合实时边缘AI场景

存算一体(PIM)已成为缓解边缘AI加速器冯·诺依曼瓶颈的主流方案。然而,现有数字PIM技术因占用大体积位单元和高密度加法树,导致计算密度低,制约宏规模扩展与能效表现。本文提出资源共享型数字存算一体单元Res-DPU,采用双端口5T SRAM锁存器与共享2T AND计算逻辑,将每比特乘法的晶体管成本降至5.25T,PIM阵列晶体管数相比最先进方案减少高达56%。同时引入新型低晶体管数二维交错加法树(TRAIT),使用FA-7T与PG-FA-26T结构,使加法树功耗降低21.35%,相比传统28T RCA设计能效提升59%。提出周期控制的迭代近似-精确乘法(CIA2M)方法,在无需纠错电路的前提下实现运行时精度-延迟动态调节。基于TSMC 65nm CMOS工艺的16KB REP-DPIM宏单元,实测达到0.43 TOPS吞吐量与87.22 TOPS/W能效,针对CIFAR-10上的ResNet-18或VGG-16模型,即使经过30%剪枝仍保持96.85%的量化质量(QoR)。研究成果为高可扩展、高能效的实时边缘AI加速器提供了可靠模块解决方案。

原文摘要 · Abstract (English)

Processing-in-memory (PIM) has emerged as the go to solution for addressing the von Neumann bottleneck in edge AI accelerators. However, state-of-the-art (SoTA) digital PIM approaches suffer from low compute density, primarily due to the use of bulky bit cells and transistor-heavy adder trees, which impose limitations on macro scalability and energy efficiency. This work introduces Res-DPU, a resource-shared digital PIM unit, with a dual-port 5T SRAM latch and shared 2T AND compute logic. This reflects the per-bit multiplication cost to just 5.25T and reduced the transistor count of the PIM array by up to 56% over the SoTA works. Furthermore, a Transistor-Reduced 2D Interspersed Adder Tree (TRAIT) with FA-7T and PG-FA-26T helps reduce the power consumption of the adder tree by up to 21.35% and leads to improved energy efficiency by 59% compared to conventional 28T RCA designs. We propose a Cycle-controlled Iterative Approximate-Accurate Multiplication (CIA2M) approach, enabling run-time accuracy-latency trade-offs without requiring error-correction circuitry. The 16 KB REP-DPIM macro achieves 0.43 TOPS throughput and 87.22 TOPS/W energy efficiency in TSMC 65nm CMOS, with 96.85% QoR for ResNet-18 or VGG-16 on CIFAR-10, including 30% pruning. The proposed results establish a Res-DPU module for highly scalable and energy-efficient real-time edge AI accelerators.

存算一体边缘计算能效优化低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。