解决模拟存内计算中电阻元件非理想响应对训练的影响
Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions

- 提出残差学习算法应对非线性响应带来的训练偏差
- 理论证明算法能精确收敛至临界点
- 适用于多种硬件缺陷,适合芯片级训练研究者
随着大规模视觉或语言模型训练与部署的经济和环境成本急剧上升,模拟存内计算(AIMC)成为一种有前景的节能解决方案。然而,其训练动态仍缺乏深入探索。在AIMC硬件中,可调权重由电阻元件的电导表示,并通过连续电脉冲更新。尽管每脉冲电导变化为常数,但实际中变化受非对称、非线性响应函数影响,导致非理想训练动态。本文为基于梯度的AIMC训练提供了理论基础,证明非对称响应函数会隐式施加目标函数惩罚。为此,我们提出残差学习算法,通过求解双层优化问题,理论上可精确收敛至临界点。该方法还可扩展至处理其他硬件缺陷,如响应粒度有限。本文首次系统研究一类通用非理想响应函数的影响,仿真验证了理论结论。
原文摘要 · Abstract (English)
As the economic and environmental costs of training and deploying large vision or language models increase dramatically, analog in-memory computing (AIMC) emerges as a promising energy-efficient solution. However, the training perspective, especially its training dynamic, is underexplored. In AIMC hardware, the trainable weights are represented by the conductance of resistive elements and updated using consecutive electrical pulses. While the conductance changes by a constant in response to each pulse, in reality, the change is scaled by asymmetric and non-linear response functions, leading to a non-ideal training dynamic. This paper provides a theoretical foundation for gradient-based training on AIMC hardware with non-ideal response functions. We demonstrate that asymmetric response functions negatively impact Analog SGD by imposing an implicit penalty on the objective. To overcome the issue, we propose Residual Learning algorithm, which provably converges exactly to a critical point by solving a bilevel optimization problem. We demonstrate that the proposed method can be extended to address other hardware imperfections, such as limited response granularity. As we know, it is the first paper to investigate the impact of a class of generic non-ideal response functions. The conclusion is supported by simulations validating our theoretical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。