arXiv:2506.00004cs.ARcs.AI2025-06被引 1

为模拟存内计算设计快速精准的电路建模方法,提升AI芯片精度。

Rapid yet accurate Tile-circuit and device modeling for Analog In-Memory Computing

  • 构建含电流压降与ADC量化效应的数学模型,快速预测运算结果。
  • 实测验证纳米级读取噪声特性,发现其在纳秒级波动显著。
  • 提出硬件感知微调策略,但传统高斯噪声对非线性压降无效。

模拟存内计算(AIMC)可使深度学习能耗降低数个数量级。然而,执行矩阵-向量乘法(MVM)的模拟‘单元’中的器件与电路非理想性会降低神经网络任务准确率。本文量化了低层失真与噪声的影响,建立了一个映射至模拟单元的乘累加(MAC)操作数学模型,完整捕捉瞬时电流压降(最主要的电路非理想性)和ADC量化效应,其预测速度远超传统精确电路仿真,且精度更高。基于实验测量,推导并匹配了纳秒级的PCM读取噪声统计模型。将这些器件(统计)与电路(确定性)效应整合进PyTorch框架,评估了BERT与ALBERT Transformer网络的精度影响。结果表明,使用简单高斯噪声进行硬件感知微调可缓解ADC量化与PCM读取噪声影响,但对压降效果不佳。这是因为压降虽为确定性,却具非线性、随时间积分窗口显著变化,且依赖于同时输入模拟单元的所有激励。训练中简单高斯噪声无法有效预训练模型应对推理时的压降,表明未来大模型部署于AIMC硬件需引入更复杂训练方法,如本文提出的单元电路模型。

原文摘要 · Abstract (English)

Analog In-Memory Compute (AIMC) can improve the energy efficiency of Deep Learning by orders of magnitude. Yet analog-domain device and circuit non-idealities -- within the analog ``Tiles'' performing Matrix-Vector Multiply (MVM) operations -- can degrade neural-network task accuracy. We quantify the impact of low-level distortions and noise, and develop a mathematical model for Multiply-ACcumulate (MAC) operations mapped to analog tiles. Instantaneous-current IR-drop (the most significant circuit non-ideality), and ADC quantization effects are fully captured by this model, which can predict MVM tile-outputs both rapidly and accurately, as compared to much slower rigorous circuit simulations. A statistical model of PCM read noise at nanosecond timescales is derived from -- and matched against -- experimental measurements. We integrate these (statistical) device and (deterministic) circuit effects into a PyTorch-based framework to assess the accuracy impact on the BERT and ALBERT Transformer networks. We show that hardware-aware fine-tuning using simple Gaussian noise provides resilience against ADC quantization and PCM read noise effects, but is less effective against IR-drop. This is because IR-drop -- although deterministic -- is non-linear, is changing significantly during the time-integration window, and is ultimately dependent on all the excitations being introduced in parallel into the analog tile. The apparent inability of simple Gaussian noise applied during training to properly prepare a DNN network for IR-drop during inference implies that more complex training approaches -- incorporating advances such as the Tile-circuit model introduced here -- will be critical for resilient deployment of large neural networks onto AIMC hardware.

存内计算神经网络硬件建模精度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。