用新方法让物理神经网络更准、更稳、能迁移。
Perfecting Imperfect Physical Neural Networks with Transferable Robustness using Sharpness-Aware Training
- 用损失曲面几何优化训练,解决模型不准问题。
- 离线训练精度超在线训练,且抗制造差异。
- 部署后自动抗温度漂移,无需重新训练。
人工智能模型在科学与工程中至关重要,但当前进展已逼近传统数字硬件的极限。为突破此限制,基于物理介质计算的物理神经网络(PNNs)受到关注。然而,其有效训练方法仍是一大挑战:现有离线训练受建模不精确制约,在线训练则生成设备专属模型,难以跨设备迁移;且部署后受热漂移、对准误差等扰动影响,模型失效需重新训练。本文提出一种名为尖锐度感知训练(Sharpness-Aware Training, SAT)的新技术,创新性地利用损失曲面几何特性,解决物理系统训练难题。SAT支持高效反向传播,即使在建模不精确时仍能实现高精度训练。离线训练的PNNs经SAT优化后,性能超越在线训练结果,且具备跨设备可迁移性。此外,SAT显著增强部署后鲁棒性,使PNN在扰动下持续准确运行而无需再训练。我们在三类PNN上验证了SAT的普适性,证明其适用于模型显式或隐式情形。该工作为模拟计算提供变革性、高效的训练方案,推动真实世界部署。
原文摘要 · Abstract (English)
AI models are essential in science and engineering, but recent advances are pushing the limits of traditional digital hardware. To address these limitations, physical neural networks (PNNs), which use physical substrates for computation, have gained increasing attention. However, developing effective training methods for PNNs remains a significant challenge. Current approaches, regardless of offline and online training, suffer from significant accuracy loss. Offline training is hindered by imprecise modeling, while online training yields device-specific models that can't be transferred to other devices due to manufacturing variances. Both methods face challenges from perturbations after deployment, such as thermal drift or alignment errors, which make trained models invalid and require retraining. Here, we address the challenges with both offline and online training through a novel technique called Sharpness-Aware Training (SAT), where we innovatively leverage the geometry of the loss landscape to tackle the problems in training physical systems. SAT enables accurate training using efficient backpropagation algorithms, even with imprecise models. PNNs trained by SAT offline even outperform those trained online, despite modeling and fabrication errors. SAT also overcomes online training limitations by enabling reliable transfer of models between devices. Finally, SAT is highly resilient to perturbations after deployment, allowing PNNs to continuously operate accurately under perturbations without retraining. We demonstrate SAT across three types of PNNs, showing it is universally applicable, regardless of whether the models are explicitly known. This work offers a transformative, efficient approach to training PNNs, addressing critical challenges in analog computing and enabling real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。