arXiv:2409.03306cs.LG2024-09NeurIPS被引 6

让模拟电路块像数字电路一样训练,实现能效突破。

Towards training digitally-tied analog blocks via hybrid gradient computation

  • 提出混合模型ff-EBMs,打通数字与模拟计算的梯度通路。
  • 在ImageNet32上达成46%准确率,刷新等效传播算法新纪录。
  • 适合关注低功耗AI训练和硬件协同设计的研究者。

标准数字电子领域的能效已趋于瓶颈,亟需新型硬件、模型与算法降低AI训练成本。基于能量的模拟电路与等效传播(EP)算法结合,构成一种有前景的梯度优化新范式。然而现有模拟加速器通常仍依赖数字电路完成非权值静态操作、补偿模拟器件缺陷,并复用已有数字加速器。这种异构硬件架构需要新的理论建模单元。本文提出前馈耦合的能量模型(ff-EBMs),融合前馈与能量型模块,分别对应数字与模拟电路。我们推导出一种新算法,通过反向传播与等效传播分别处理前馈与能量部分,实现对ff-EBMs的端到端梯度计算,使EP可应用于更灵活、更真实的架构。实验表明,在ff-EBMs中使用深度霍普菲尔德网络(DHNs)作为能量块时,标准DHN可任意均匀拆分且性能不变;在ImageNet32上训练取得46%的top-1准确率,为EP领域当前最佳表现。该方法提供了一条可扩展、可增量地将自训练模拟计算单元集成至现有数字加速器的系统性路径。

原文摘要 · Abstract (English)

Power efficiency is plateauing in the standard digital electronics realm such that novel hardware, models, and algorithms are needed to reduce the costs of AI training. The combination of energy-based analog circuits and the Equilibrium Propagation (EP) algorithm constitutes one compelling alternative compute paradigm for gradient-based optimization of neural nets. Existing analog hardware accelerators, however, typically incorporate digital circuitry to sustain auxiliary non-weight-stationary operations, mitigate analog device imperfections, and leverage existing digital accelerators.This heterogeneous hardware approach calls for a new theoretical model building block. In this work, we introduce Feedforward-tied Energy-based Models (ff-EBMs), a hybrid model comprising feedforward and energy-based blocks accounting for digital and analog circuits. We derive a novel algorithm to compute gradients end-to-end in ff-EBMs by backpropagating and "eq-propagating" through feedforward and energy-based parts respectively, enabling EP to be applied to much more flexible and realistic architectures. We experimentally demonstrate the effectiveness of the proposed approach on ff-EBMs where Deep Hopfield Networks (DHNs) are used as energy-based blocks. We first show that a standard DHN can be arbitrarily split into any uniform size while maintaining performance. We then train ff-EBMs on ImageNet32 where we establish new SOTA performance in the EP literature (46 top-1 %). Our approach offers a principled, scalable, and incremental roadmap to gradually integrate self-trainable analog computational primitives into existing digital accelerators.

模拟计算能量模型等效传播能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。