无需反向传播,用光芯片加速物理神经网络训练,实现实时更新。
Scalable Back-Propagation-Free Training of Optical Physics-Informed Neural Networks
- 用稀疏网格斯坦导数估计替代反向传播,避免训练瓶颈。
- 通过张量分解实现低阶优化,提升大规模训练的可扩展性。
- 光子张量核心设计使芯片面积大幅缩减,支持实时训练。
物理智能与数字孪生系统需对机器人、自动驾驶、芯片等工程系统进行快速反复评估,以支持近乎实时的决策。现有高性能计算设备(如GPU)上训练物理信息神经网络(PINNs)耗时仍远超实时需求。光子计算凭借超高速运算潜力,有望填补这一延迟鸿沟。但光子器件缺乏内存且尺寸大,难以在芯片上训练大型PINNs。本文提出一种完全无反向传播(BP-free)且高度可扩展的框架,在硅光平台训练真实规模的PINNs。方法包括:(1) 稀疏网格斯坦导数估计器,避免损失函数中的反向传播;(2) 基于张量-列车分解的降维零阶优化,提升可扩展性与收敛性;(3) 可扩展的片上光子PINN训练加速器设计,采用光子张量核心。我们在低维与高维偏微分方程(PDE)基准上验证了数值方法的有效性。基于真实器件参数的预硅仿真进一步证明,该光子加速器在实时训练和芯片面积显著缩减方面具有明显优势。
原文摘要 · Abstract (English)
Physics intelligence and digital twins often require rapid and repeated performance evaluation of various engineering systems (e.g. robots, autonomous vehicles, semiconductor chips) to enable (almost) real-time actions or decision making. This has motivated the development of accelerated partial differential equation (PDE) solvers, in resource-constrained scenarios if the PDE solvers are to be deployed on the edge. Physics-informed neural networks (PINNs) have shown promise in solving high-dimensional PDEs, but the training time on state-of-the-art digital hardware (e.g., GPUs) is still orders-of-magnitude longer than the latency required for enabling real-time decision making. Photonic computing offers a potential solution to address this huge latency gap because of its ultra-high operation speed. However, the lack of photonic memory and the large device sizes prevent training real-size PINNs on photonic chips. This paper proposes a completely back-propagation-free (BP-free) and highly salable framework for training real-size PINNs on silicon photonic platforms. Our approach involves three key innovations: (1) a sparse-grid Stein derivative estimator to avoid the BP in the loss evaluation of a PINN, (2) a dimension-reduced zeroth-order optimization via tensor-train decomposition to achieve better scalability and convergence in BP-free training, and (3) a scalable on-chip photonic PINN training accelerator design using photonic tensor cores. We validate our numerical methods on both low- and high-dimensional PDE benchmarks. Through pre-silicon simulation based on real device parameters, we further demonstrate the significant performance benefit (e.g., real-time training, huge chip area reduction) of our photonic accelerator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。