用65纳米芯片实现低功耗图像分类,每帧仅耗8.6纳焦。
An All-digital 8.6-nJ/Frame 65-nm Tsetlin Machine Image Classification Accelerator
- 全数字架构,基于命题逻辑的Tsetlin机并行加速
- 每秒处理60.3千帧,28×28图像分类延迟25.4微秒
- 适用于边缘设备的高能效推理,适合低功耗场景
我们提出一种基于命题逻辑的可编程机器学习加速器芯片,采用Tsetlin机(TM)原理,通过子模式识别表达式(即规则)进行图像分类。该加速器实现合并版Tsetlin机与卷积结合,对28×28像素的布尔化图像进行10类分类,使用128条规则,采用高度并行架构。所有规则权重和Tsetlin自动机动作信号保留在寄存器中,实现快速规则评估。芯片采用65纳米低漏电CMOS工艺,功耗面积为2.7 mm²。在27.8 MHz时钟频率下,每秒完成60.3千次分类,每帧仅耗8.6纳焦。单帧分类延迟为25.4微秒(含系统开销)。在MNIST、Fashion-MNIST和Kuzushiji-MNIST数据集上,测试准确率分别为97.42%、84.54%和82.55%,与软件模型一致。
原文摘要 · Abstract (English)
We present an all-digital programmable machine learning accelerator chip for image classification, underpinning on the Tsetlin machine (TM) principles. The TM is an emerging machine learning algorithm founded on propositional logic, utilizing sub-pattern recognition expressions called clauses. The accelerator implements the coalesced TM version with convolution, and classifies booleanized images of 28$\times$28 pixels with 10 categories. A configuration with 128 clauses is used in a highly parallel architecture. Fast clause evaluation is achieved by keeping all clause weights and Tsetlin automata (TA) action signals in registers. The chip is implemented in a 65 nm low-leakage CMOS technology, and occupies an active area of 2.7 mm$^2$. At a clock frequency of 27.8 MHz, the accelerator achieves 60.3k classifications per second, and consumes 8.6 nJ per classification. This demonstrates the energy-efficiency of the TM, which was the main motivation for developing this chip. The latency for classifying a single image is 25.4 $μ$s which includes system timing overhead. The accelerator achieves 97.42%, 84.54% and 82.55% test accuracies for the datasets MNIST, Fashion-MNIST and Kuzushiji-MNIST, respectively, matching the TM software models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。