arXiv:2412.05327cs.ARcs.AI2024-12

用Y-Flash内存芯片实现高效推理,专为逻辑型机器学习设计。

IMPACT:InMemory ComPuting Architecture Based on Y-FlAsh Technology for Coalesced Tsetlin Machine Inference

  • 基于Y-Flash内存阵列构建双交叉条结构,实现并行化推理
  • 在MNIST上达到96.3%准确率,能效比领先传统方案2倍以上
  • 适合低功耗、高能效的现代机器学习推理场景

随着对大规模数据处理需求的增长,传统冯·诺依曼架构已难以满足日益提升的数据带宽要求。存内计算(IMC)通过在微架构层面实现数据的分布式存储与处理,显著降低延迟和能耗,成为新兴解决方案。本文提出基于180 nm CMOS工艺制造的新型存储器件Y-Flash的存内计算架构IMPACT,用于协同式刘易斯机(Coalesced Tsetlin Machine, CoTM)推理。CoTM利用刘易斯自动机(TA)在并行子句中随机生成布尔特征选择,借助两组计算交叉条阵列分别存储TA状态与权重。在MNIST数据集上的验证表明,该架构实现96.3%的识别准确率。相比基于忆阻器的CNN,能效提升2.23倍;相比基于NOR-Flash的类脑系统,提升2.46倍;相比基于相变存储器的深度神经网络,提升2.06倍,适用于现代机器学习推理应用。

原文摘要 · Abstract (English)

The increasing demand for processing large volumes of data for machine learning models has pushed data bandwidth requirements beyond the capability of traditional von Neumann architecture. In-memory computing (IMC) has recently emerged as a promising solution to address this gap by enabling distributed data storage and processing at the micro-architectural level, significantly reducing both latency and energy. In this paper, we present the IMPACT: InMemory ComPuting Architecture Based on Y-FlAsh Technology for Coalesced Tsetlin Machine Inference, underpinned on a cutting-edge memory device, Y-Flash, fabricated on a 180 nm CMOS process. Y-Flash devices have recently been demonstrated for digital and analog memory applications, offering high yield, non-volatility, and low power consumption. The IMPACT leverages the Y-Flash array to implement the inference of a novel machine learning algorithm: coalesced Tsetlin machine (CoTM) based on propositional logic. CoTM utilizes Tsetlin automata (TA) to create Boolean feature selections stochastically across parallel clauses. The IMPACT is organized into two computational crossbars for storing the TA and weights. Through validation on the MNIST dataset, IMPACT achieved 96.3% accuracy. The IMPACT demonstrated improvements in energy efficiency, e.g., 2.23X over CNN-based ReRAM, 2.46X over Neuromorphic using NOR-Flash, and 2.06X over DNN-based PCM, suited for modern ML inference applications.

存内计算Y-Flash逻辑推理能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。