arXiv:2501.13181cs.ARcs.AI2025-01

用对数域模拟电路实现低功耗神经网络训练加速。

Learning in Log-Domain: Subthreshold Analog AI Accelerator Based on Stochastic Gradient Descent

  • 在亚阈值区用对数域电路实现随机梯度下降训练。
  • 误差低于0.87%,8位精度下仍保持高逼近性。
  • 适合需要边缘端在线训练的低功耗场景。

AI模型的快速普及和边缘部署需求,推动了高性能、低功耗AI硬件的发展。本文提出一种面向基于L2正则化的随机梯度下降(SGDr)训练任务的新型模拟加速器架构。该架构采用亚阈值MOS器件的对数域电路,并结合易失性存储单元。我们建立了连续时间域求解SGDr的数学框架,并详细说明了学习方程到对数域电路的映射方法。通过模拟域运算和弱反型工作模式,与数字实现相比,该设计显著降低了晶体管面积和功耗。实验表明,该架构能紧密逼近理想行为,均方误差低于0.87%,精度可达8位。此外,该架构支持广泛的超参数配置。本工作为具备片上训练能力的低功耗模拟AI硬件开辟了新路径。

原文摘要 · Abstract (English)

The rapid proliferation of AI models, coupled with growing demand for edge deployment, necessitates the development of AI hardware that is both high-performance and energy-efficient. In this paper, we propose a novel analog accelerator architecture designed for AI/ML training workloads using stochastic gradient descent with L2 regularization (SGDr). The architecture leverages log-domain circuits in subthreshold MOS and incorporates volatile memory. We establish a mathematical framework for solving SGDr in the continuous time domain and detail the mapping of SGDr learning equations to log-domain circuits. By operating in the analog domain and utilizing weak inversion, the proposed design achieves significant reductions in transistor area and power consumption compared to digital implementations. Experimental results demonstrate that the architecture closely approximates ideal behavior, with a mean square error below 0.87% and precision as low as 8 bits. Furthermore, the architecture supports a wide range of hyperparameters. This work paves the way for energy-efficient analog AI hardware with on-chip training capabilities.

模拟计算低功耗边缘训练亚阈值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。