用热力学硬件加速二阶优化,显著提升大规模模型训练速度。
Scalable Thermodynamic Second-order Optimization
- 设计可扩展算法,利用热力学计算机实现K-FAC二阶优化
- 量化噪声下仍保持二阶优化优势,n越大越明显
- 适合大规模视觉与图神经网络训练场景
许多硬件方案致力于加速AI推理,但对训练加速关注较少,尽管快速训练对社会影响巨大。基于物理的计算机(如热力学计算机)能高效解决AI训练中的核心计算问题。在数字硬件上因矩阵求逆代价高而难以实现的优化器,可通过物理硬件解锁。本文提出一种可扩展算法,利用热力学计算机加速流行的二阶优化器Kronecker-factored approximate curvature(K-FAC)。渐近复杂度分析表明,随着每层神经元数n增加,该算法优势愈发显著。数值实验显示,即使存在显著量化噪声,二阶优化的优势仍可保持。基于真实硬件特性预测,大规模视觉与图任务将获得显著加速。
原文摘要 · Abstract (English)
Many hardware proposals have aimed to accelerate inference in AI workloads. Less attention has been paid to hardware acceleration of training, despite the enormous societal impact of rapid training of AI models. Physics-based computers, such as thermodynamic computers, offer an efficient means to solve key primitives in AI training algorithms. Optimizers that normally would be computationally out-of-reach (e.g., due to expensive matrix inversions) on digital hardware could be unlocked with physics-based hardware. In this work, we propose a scalable algorithm for employing thermodynamic computers to accelerate a popular second-order optimizer called Kronecker-factored approximate curvature (K-FAC). Our asymptotic complexity analysis predicts increasing advantage with our algorithm as $n$, the number of neurons per layer, increases. Numerical experiments show that even under significant quantization noise, the benefits of second-order optimization can be preserved. Finally, we predict substantial speedups for large-scale vision and graph problems based on realistic hardware characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。