光子卷积神经网络实现高能效图像分类,精度超94%。
Photonic convolutional neural network with pre-trained in situ training

- 全光域完成卷积、池化、非线性激活等操作,无需电子转换。
- 在MNIST上达94.49%准确率,比电子芯片能效高220至330倍。
- 支持预训练与硬件微调,对制造误差有天然鲁棒性。
卷积神经网络(CNN)已重塑图像处理,但基于电子芯片的实现存在能耗和推理延迟的根本瓶颈。这推动了超越互补金属氧化物半导体(CMOS)芯片的新型硬件架构探索。光学系统可在光速下以极低能耗执行线性矩阵运算,适用于CNN加速。然而,构建一个能同时完成线性与非线性运算且高效训练的全相干光子CNN仍是开放挑战。本文提出一种全光子卷积神经网络(PCNN),在光学域完成图像分类,包含卷积、最大池化、非线性激活及全连接层。该网络在马赫-曾德尔干涉仪(MZI)阵列、多模干涉仪(MMI)树及微环谐振器非线性模块上实现,于MNIST数据集达到94.49%准确率。通过数学精确的可微分数字孪生模型进行离线预训练,达到97.45%数字精度;训练相位直接映射至光子硬件,并通过仅需两次前向传播的无梯度算法进行优化。该架构对传播损耗、MZI插入损耗、制造偏差及热串扰具有内在鲁棒性。底层功耗分析显示,芯片静态功耗为10.83 W,推理延迟为843 ns,单图推理能效比当前最先进电子GPU高220至330倍。
原文摘要 · Abstract (English)
Convolutional neural networks (CNNs) have transformed image processing, but the energy consumption and inference latency of electronic based implementations remain fundamental bottlenecks. These limitations have motivated the search for alternative hardware architectures beyond Complementary metal-oxide-semiconductor (CMOS) chips. Optical systems can perform linear matrix operations at the speed of light with extremely low energy dissipation, making them attractive for CNN acceleration. However, building a fully coherent photonic CNN that performs both linear and nonlinear operations and training it efficiently remains an open challenge. Here we present a fully photonic convolutional neural network (PCNN) that executes image classification in the optical domain, including convolution, max-pooling, nonlinear activation, and fully connected layers. The network achieves 94.49 percent accuracy on the MNIST dataset distributed across Mach Zehnder Interferometer (MZI) meshes, weighted Multimode Interferometer (MMI) trees, and a microring resonator based nonlinearity. A mathematically exact differentiable digital twin, enables backpropagation for ex situ pre training, reaches 97.45 percent digital accuracy. Trained phases are transferred one-to-one to the photonic hardware and refined via a gradient free algorithm that estimates the full gradient with only two forward passes. The architecture exhibits inherent robustness to non idealities, under the compound effect of propagation loss, MZI insertion loss, fabrication disorder, and thermal crosstalk. A bottom-up power analysis yields 10.83 W static chip consumption and 843 ns inference latency, translating to 220 to 330 times greater energy efficiency than state of the art electronic GPUs for single-image inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。