arXiv:2606.07896physics.opticscs.CV2026-06

提出可微分体传播算法,让可见光全息神经网络设计更贴近实际制造。

Beyond the Thin-Layer Limit: Differentiable Volumetric Training for Visible-Range Diffractive Neural Networks

论文配图:Beyond the Thin-Layer Limit: Differentiable Volumetric Training for Visible-Range Diffractive Neural Networks
图 1 · 摘自论文原文
  • 用可微分体传播建模光在厚层中的衍射与相位积累
  • 在多个数据集上将设备实现准确率从50%提升至90%
  • 适合做光学神经网络设计的科研人员和工程师

全息深度神经网络(D2NN)有望为机器视觉提供微型化、低功耗、光速的光学前端,但现有成熟成果多限于太赫兹波段,依赖毫米级神经元。将D2NN拓展至可见光波段(主流视觉系统工作范围)长期被认为受限于纳米级神经元的难以制造;然而即便近期制造障碍已消除,可见光波段的D2NN性能仍远不及太赫兹波段。我们发现真正瓶颈在于几乎所有D2NN训练中采用的薄层近似——将每个衍射层视为无限薄掩膜。这一问题并非由短波长导致,而是因为可见光常用低折射率材料(n≈1.3-1.5),需足够厚度的浮雕结构,从而引发显著的层内衍射与相位累积。为此,我们提出可微分束传播(∂BPM)层,将每个单元建模为有限厚度体积,在训练中模拟光通过过程,保持制造兼容的高度图端到端可训练,无需全程全波仿真。在MNIST、Fashion-MNIST和CIFAR-100分类与成像任务中,∂BPM训练显著降低设计到器件的失配,全波FDTD验证后分类准确率从50%提升至90%,且无需重新优化。∂BPM层为高效光学神经网络优化与制造一致的衍射设计之间提供了可扩展、物理感知的桥梁。

原文摘要 · Abstract (English)

Diffractive deep neural networks (D2NNs) promise miniaturized, power-efficient, light-speed optical front-ends for machine vision, yet the most mature demonstrations remain in the terahertz regime, built from readily fabricated millimeter-scale neurons. Translating D2NNs to the visible range, where nearly all vision pipelines operate, was long blamed on the difficulty of fabricating nanoscale neurons; but even after recent advances removed that barrier, visible-range D2NNs matching their terahertz counterparts remain out of reach. We identify the true obstacle as the thin-layer approximation underlying nearly all D2NN training, which treats each diffractive layer as an infinitely thin mask. It fails not because of the short wavelength, as is commonly assumed, but because the low-refractive-index materials (n approximately 1.3-1.5) used at visible wavelengths require relief structures thick enough that intra-layer diffraction and phase accumulation become significant. To overcome this, we introduce a differentiable beam-propagation ($\partial$BPM) layer that models each element as a finite-thickness volume and propagates light through it during training, keeping the fabrication-compatible height map end-to-end trainable without full-wave simulation in the loop. Across MNIST, Fashion-MNIST, and CIFAR-100 classification and imaging, $\partial$BPM training substantially reduces the design-to-device mismatch, and full-wave FDTD validation raises classification accuracy from 50% to 90% without re-optimization. The $\partial$BPM layer thus offers a scalable, physics-aware bridge between efficient optical neural-network optimization and fabrication-consistent diffractive design.

光学神经网络可微分物理全息计算制造一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。