用反向传播训练热力学神经网络,在低功耗硬件上实现高精度图像分类。
Scaling Up Thermodynamic AI Models

- 基于吉布斯采样设计可扩展的反向传播训练算法
- 在CIFAR-10和CIFAR-100上分别达到94.9%和76.0%准确率
- 揭示推理成本与精度的数学关系,指导硬件优化
基于伊辛模型的热力学计算设备在低功耗人工智能推理与边缘计算中展现出巨大潜力,但针对此类硬件的大规模模型训练方法仍受限。已有理论表明,高温吉布斯采样的时间平均行为可实现前馈神经网络推理。本文将该理论对应转化为一种可扩展的纯反向传播训练算法,用于在伊辛机硬件上进行深度卷积网络的热力学推理训练。所训练的图像分类模型在二元吉布斯采样下于CIFAR-10上达到94.9%准确率,在CIFAR-100上达到76.0%。随后,我们建立并实验验证了推理成本与准确率及自相关时间之间的数学理论关系。进一步推导出渐近结果,表明推理成本受性能与控制参数间明确权衡的约束,并提出计算最优推理调度的算法。最后讨论了对硬件开发及高温热力学人工智能未来的影响。
原文摘要 · Abstract (English)
Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited. Prior theory shows that the time-averaged behavior of high-temperature Gibbs-sampled Ising systems can implement feed-forward neural inference. We turn this theoretical correspondence into a scalable and purely backpropagation-based algorithm for training deep convolutional networks for thermodynamic inference on Ising machine hardware. Our image classification models achieve accuracies of 94.9% on CIFAR-10 and 76.0% on CIFAR-100 under binary Gibbs sampling. We then develop and experimentally validate a mathematical theory relating inference cost to accuracy and controlling autocorrelation times. Subsequently, we calculate asymptotic results showing that inference cost is bounded by a well-controlled tradeoff with performance and exhibit algorithms for computing optimal inference schedules. Finally, we discuss implications for hardware development and the future of high-temperature thermodynamic AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。