提出新型进化剪枝算法,高效训练大规模光神经网络并提升抗噪能力。
Efficient training for large-scale optical neural network using an evolutionary strategy and attention pruning
- 用进化策略与注意力剪枝联合优化光网络参数
- 剪掉60%-80%参数,性能损失仅3.3%-4.7%
- 在噪声环境下仍优于传统方法,适合大规模部署
基于MZI的块光神经网络(BONNs)可实现大规模模型,但现有训练算法鲁棒性不足,且参数量大导致计算和功耗高。本文提出一种片上协方差矩阵自适应进化策略与注意力剪枝(CAP)算法,通过剪枝矩阵块并直接优化种群个体,显著降低复杂度。实验表明,对MNIST和Fashion-MNIST数据集,CAP分别可剪枝60%和80%参数,性能仅下降3.289%和4.693%。在存在相位器动态噪声(标准差0.5)的劣质芯片上,其性能下降分别为22.327%和24.019%,显著优于先前块伴随训练算法(43.963%、41.074%)和标准进化策略(25.757%、32.871%)。当剪枝60%参数时,实验中简化MNIST数据集准确率达88.5%,接近无噪声模拟结果(92.1%)。仿真与实验证明,仅使用内部相位移器的MZI构建BONNs可有效减少系统面积与可训练参数。该算法在更大规模模型与更复杂任务中展现出巨大潜力。
原文摘要 · Abstract (English)
MZI-based block optical neural networks (BONNs), which can achieve large-scale network models, have increasingly drawn attentions. However, the robustness of the current training algorithm is not high enough. Moreover, large-scale BONNs usually contain numerous trainable parameters, resulting in expensive computation and power consumption. In this article, by pruning matrix blocks and directly optimizing the individuals in population, we propose an on-chip covariance matrix adaptation evolution strategy and attention-based pruning (CAP) algorithm for large-scale BONNs. The calculated results demonstrate that the CAP algorithm can prune 60% and 80% of the parameters for MNIST and Fashion-MNIST datasets, respectively, while only degrades the performance by 3.289% and 4.693%. Considering the influence of dynamic noise in phase shifters, our proposed CAP algorithm (performance degradation of 22.327% for MNIST dataset and 24.019% for Fashion-MNIST dataset utilizing a poor fabricated chip and electrical control with a standard deviation of 0.5) exhibits strongest robustness compared with both our previously reported block adjoint training algorithm (43.963% and 41.074%) and the covariance matrix adaptation evolution strategy (25.757% and 32.871%), respectively. Moreover, when 60% of the parameters are pruned, the CAP algorithm realizes 88.5% accuracy in experiment for the simplified MNIST dataset, which is similar to the simulation result without noise (92.1%). Additionally, we simulationally and experimentally demonstrate that using MZIs with only internal phase shifters to construct BONNs is an efficient way to reduce both the system area and the required trainable parameters. Notably, our proposed CAP algorithm show excellent potential for larger-scale network models and more complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。