用模拟存内计算加速核方法,提升能效与速度。
Kernel Approximation using Analog In-Memory Computing
- 在内存中直接执行核近似计算,减少数据搬移
- 分类任务准确率下降小于1%,注意力机制误差在1%内
- 适合追求低功耗的边缘AI部署
核函数是多种机器学习算法的核心,但常带来高昂的内存与计算开销。本文提出一种适用于混合信号模拟存内计算(AIMC)架构的核近似方法。该方法通过在内存中直接执行核近似运算,缓解传统核方法的性能瓶颈。基于相变存储器的IBM HERMES项目芯片完成了硬件验证。实验表明,该方法在核化岭分类基准上准确率损失低于1%,在Transformer的核化注意力长程任务(Long Range Arena)上保持99%以上准确率。相比传统数字加速器,本方案预计能实现更高能效与更低功耗。结果表明,异构存内计算架构在提升机器学习应用效率与可扩展性方面具有潜力。
原文摘要 · Abstract (English)
Kernel functions are vital ingredients of several machine learning algorithms, but often incur significant memory and computational costs. We introduce an approach to kernel approximation in machine learning algorithms suitable for mixed-signal Analog In-Memory Computing (AIMC) architectures. Analog In-Memory Kernel Approximation addresses the performance bottlenecks of conventional kernel-based methods by executing most operations in approximate kernel methods directly in memory. The IBM HERMES Project Chip, a state-of-the-art phase-change memory based AIMC chip, is utilized for the hardware demonstration of kernel approximation. Experimental results show that our method maintains high accuracy, with less than a 1% drop in kernel-based ridge classification benchmarks and within 1% accuracy on the Long Range Arena benchmark for kernelized attention in Transformer neural networks. Compared to traditional digital accelerators, our approach is estimated to deliver superior energy efficiency and lower power consumption. These findings highlight the potential of heterogeneous AIMC architectures to enhance the efficiency and scalability of machine learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。