arXiv:2511.17265cs.ARcs.AI2025-11中稿 · publication in the…被引 2

DISCA用压缩的弯金字塔格式实现高效低功耗矩阵乘法,适合边缘AI硬件。

DISCA: A Digital In-memory Stochastic Computing Architecture Using A Compressed Bent-Pyramid Format

  • 采用压缩弯金字塔格式的数字存内随机计算架构
  • 500MHz下能效达3.59TOPS/W(180nm工艺)
  • 兼顾模拟计算简洁性与数字系统可靠性,适合边缘设备

当前人工智能在医疗、机器人、自动驾驶、安防等领域广泛应用,大规模模型依赖海量参数和密集矩阵乘法。传统冯·诺依曼架构受限于内存墙和摩尔定律终结,而边缘设备如机器人、无人机对硬件预算要求严苛。尽管存内计算有望突破内存墙,但模拟与数字方案均因设计限制导致性能衰减。本文提出新型数字存内随机计算架构DISCA,采用压缩的准随机弯金字塔数据格式。DISCA兼具模拟计算的简单性和数字系统的可扩展性、生产率与可靠性。后版图建模结果显示,在500 MHz下使用商用180 nm CMOS工艺,能效达3.59TOPS/W。相比同类架构,该方案在规模扩展后可实现矩阵乘法能效量级提升。

原文摘要 · Abstract (English)

Nowadays, we are witnessing an Artificial Intelligence revolution that dominates the technology landscape in various application domains, such as healthcare, robotics, automotive, security, and defense. Massive-scale AI models, which mimic the human brain's functionality, typically feature millions and even billions of parameters through data-intensive matrix multiplication tasks. While conventional Von-Neumann architectures struggle with the memory wall and the end of Moore's Law, these AI applications are migrating rapidly towards the edge, such as in robotics and unmanned aerial vehicles for surveillance, thereby adding more constraints to the hardware budget of AI architectures at the edge. Although in-memory computing has been proposed as a promising solution for the memory wall, both analog and digital in-memory computing architectures suffer from substantial degradation of the proposed benefits due to various design limitations. We propose a new digital in-memory stochastic computing architecture, DISCA, utilizing a compressed version of the quasi-stochastic Bent-Pyramid data format. DISCA inherits the same computational simplicity of analog computing, while preserving the same scalability, productivity, and reliability of digital systems. Post-layout modeling results of DISCA show an energy efficiency of 3.59TOPS/W per bit at 500 MHz using a commercial 180 nm CMOS technology. Therefore, DISCA significantly improves the energy efficiency for matrix multiplication workloads by orders of magnitude if scaled and compared to its counterpart architectures.

存内计算随机计算边缘AI能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。