arXiv:2509.18121cs.ETcs.LG2025-09被引 2

用短脉冲提升神经网络训练能效,兼顾精度与速度

Energy-convergence trade off for the training of neural networks on bio-inspired hardware

  • 用实验测得的铁电忆阻器更新数据模拟训练,优化脉冲宽度
  • 20纳秒脉冲可降低每步能耗,总能耗减少但需更多迭代次数
  • 提出对称点偏移法,解决非对称更新问题,恢复精度

可穿戴与植入式设备的普及将人工智能计算推向极致边缘,要求极低功耗以实现持续运行。受大脑启发,新兴的忆阻器件有望通过消除计算与存储间的数据传输,加速神经网络训练。然而性能与能效的平衡仍是挑战。本文研究基于HfO2/ZrO2超晶格的铁电突触器件,将其实验测得的权重更新结果用于硬件感知的神经网络仿真。在20纳秒至0.2毫秒的脉冲宽度范围内,较短脉冲可降低单次更新能耗,但需更多训练轮次,整体仍能减少总能耗且不损失精度。使用普通随机梯度下降(SGD)时分类准确率低于混合精度SGD。我们分析原因并提出“对称点偏移”技术,缓解非对称更新问题,恢复模型精度。结果揭示了精度、收敛速度与能耗间的权衡关系,表明结合特定训练策略的短脉冲编程可显著提升片上学习效率。

原文摘要 · Abstract (English)

The increasing deployment of wearable sensors and implantable devices is shifting AI processing demands to the extreme edge, necessitating ultra-low power for continuous operation. Inspired by the brain, emerging memristive devices promise to accelerate neural network training by eliminating costly data transfers between compute and memory. Though, balancing performance and energy efficiency remains a challenge. We investigate ferroelectric synaptic devices based on HfO2/ZrO2 superlattices and feed their experimentally measured weight updates into hardware-aware neural network simulations. Across pulse widths from 20 ns to 0.2 ms, shorter pulses lower per-update energy but require more training epochs while still reducing total energy without sacrificing accuracy. Classification accuracy using plain stochastic gradient descent (SGD) is diminished compared to mixed-precision SGD. We analyze the causes and propose a ``symmetry point shifting'' technique, addressing asymmetric updates and restoring accuracy. These results highlight a trade-off among accuracy, convergence speed, and energy use, showing that short-pulse programming with tailored training significantly enhances on-chip learning efficiency.

神经网络训练忆阻器能效优化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。