arXiv:2509.15113cs.LG2025-09

让物理层黑盒模型可端到端训练,实现高效混合神经网络。

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers

  • 用低秩代理模型模拟物理层梯度传播
  • 零阶随机优化更新物理层参数,接近数字基线精度
  • 适合光子/类脑硬件与神经网络融合的开发者

为实现能效高、性能强的AI系统,研究者关注光子、类脑等新型计算平台,但其非可微特性使反向传播难以实施。本文提出一种混合网络端到端训练框架:通过动态低秩代理模型实现梯度穿越物理层,结合随机零阶优化更新物理层内部参数。关键创新是隐式投影分裂积分算法,在每次前向传播后以极低硬件查询代价更新轻量代理模型,避免全矩阵重构。在计算机视觉、音频分类和语言建模任务中,该方法均达到接近数字基线的精度,成功实现空间光调制器、微环谐振器、马赫-曾德干涉仪等非可微物理组件的有效集成。本工作打通了硬件感知深度学习与无梯度优化的壁垒,为可扩展端到端可训练系统提供了实用路径。

原文摘要 · Abstract (English)

The growing demand for energy-efficient, high-performance AI systems has led to increased attention on alternative computing platforms (e.g., photonic, neuromorphic) due to their potential to accelerate learning and inference. However, integrating such physical components into deep learning pipelines remains challenging, as physical devices often offer limited expressiveness, and their non-differentiable nature renders on-device backpropagation difficult or infeasible. This motivates the development of hybrid architectures that combine digital neural networks with reconfigurable physical layers, which effectively behave as black boxes. In this work, we present a framework for the end-to-end training of such hybrid networks. This framework integrates stochastic zeroth-order optimization for updating the physical layer's internal parameters with a dynamic low-rank surrogate model that enables gradient propagation through the physical layer. A key component of our approach is the implicit projector-splitting integrator algorithm, which updates the lightweight surrogate model after each forward pass with minimal hardware queries, thereby avoiding costly full matrix reconstruction. We demonstrate our method across diverse deep learning tasks, including: computer vision, audio classification, and language modeling. Notably, across all modalities, the proposed approach achieves near-digital baseline accuracy and consistently enables effective end-to-end training of hybrid models incorporating various non-differentiable physical components (spatial light modulators, microring resonators, and Mach-Zehnder interferometers). This work bridges hardware-aware deep learning and gradient-free optimization, thereby offering a practical pathway for integrating non-differentiable physical components into scalable, end-to-end trainable AI systems.

混合架构零阶优化物理层低秩模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。