arXiv:2501.07991physics.opticscs.AI2025-01被引 6

用光子非线性模块替代部分神经网络,大幅降低算力消耗。

Training Hybrid Neural Networks with Multimode Optical Nonlinearities Using Digital Twins

  • 用多模光纤的超短脉冲传播模拟非线性计算,作为固定计算模块
  • 实验达到顶尖图像分类准确率,且模拟与真实系统高度一致
  • 对实验漂移有强鲁棒性,适合低功耗边缘AI部署

训练越来越大的神经网络推动了人工智能在科学与技术中的应用,但其指数级增长带来了更高的能耗与硬件需求。将复杂的物理过程作为固定、高效的计算模块嵌入网络,可降低可训练层的复杂度。本文利用多模光纤中超短脉冲的传播,实现大规模非线性变换,作为此类计算模块。通过可微分近似的神经模型构建光学系统的数字孪生,训练算法更新该仿真器,并沿代理反向传播误差信号以优化光学模块前的层。实验结果表明,该混合架构实现了当前最优的图像分类精度和模拟保真度,同时展现出对实验漂移的出色鲁棒性。该方法通过整合低能耗物理系统,使神经网络具备可扩展性与能效优势,显著减少计算负担。

原文摘要 · Abstract (English)

The ability to train ever-larger neural networks brings artificial intelligence to the forefront of scientific and technical discoveries. However, their exponentially increasing size creates a proportionally greater demand for energy and computational hardware. Incorporating complex physical events in networks as fixed, efficient computation modules can address this demand by decreasing the complexity of trainable layers. Here, we utilize ultrashort pulse propagation in multimode fibers, which perform large-scale nonlinear transformations, for this purpose. Training the hybrid architecture is achieved through a neural model that differentiably approximates the optical system. The training algorithm updates the neural simulator and backpropagates the error signal over this proxy to optimize layers preceding the optical one. Our experimental results achieve state-of-the-art image classification accuracies and simulation fidelity. Moreover, the framework demonstrates exceptional resilience to experimental drifts. By integrating low-energy physical systems into neural networks, this approach enables scalable, energy-efficient AI models with significantly reduced computational demands.

光子计算混合架构低功耗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。