用超表面实现可任意设计的光学卷积核,加速机器视觉任务。
Metasurface-generated large and arbitrary analog convolution kernels for accelerated machine vision
- 通过频域训练法在超表面上生成任意形状的模拟卷积核。
- 在MNIST上达98.59%准确率,Fashion-MNIST和CIFAR-10模拟结果分别达92.63%和68.67%。
- 适合边缘设备部署,特别适用于对速度和功耗敏感的实时视觉系统。
在人工智能快速发展的背景下,卷积神经网络在机器视觉与医学诊断等复杂任务中至关重要。为应对传统数字卷积运算的处理速度慢、功耗高等问题,研究者提出用光学元件替代数字卷积层以加速各类机器视觉任务。然而,光学卷积核的模拟特性尚未被充分挖掘。本文提出一种空间频率域训练方法,利用光学超表面作为卷积层,生成任意形状的模拟卷积核,其感受野显著超过数字卷积核。在非相干照明条件下,通过空间复用实现了多个并行卷积核(含正负权重)的生成。实验在MNIST数据集上实现98.59%分类准确率;模拟结果显示,在加入数字层后,Fashion-MNIST和CIFAR-10数据集的准确率分别为92.63%和68.67%。该工作凸显了模拟光学卷积的独特优势,为加速机器视觉任务提供了可行路径,尤其适用于边缘设备。
原文摘要 · Abstract (English)
In the rapidly evolving field of artificial intelligence, convolutional neural networks are essential for tackling complex challenges such as machine vision and medical diagnosis. Recently, to address the challenges in processing speed and power consumption of conventional digital convolution operations, many optical components have been suggested to replace the digital convolution layer in the neural network, accelerating various machine vision tasks. Nonetheless, the analog nature of the optical convolution kernel has not been fully explored. Here, we develop a spatial frequency domain training method to create arbitrarily shaped analog convolution kernels using an optical metasurface as the convolution layer, with its receptive field largely surpassing digital convolution kernels. By employing spatial multiplexing, the multiple parallel convolution kernels with both positive and negative weights are generated under the incoherent illumination condition. We experimentally demonstrate a 98.59% classification accuracy on the MNIST dataset, with simulations showing 92.63% and 68.67% accuracy on the Fashion-MNIST and CIFAR-10 datasets with additional digital layers. This work underscores the unique advantage of analog optical convolution, offering a promising avenue to accelerate machine vision tasks, especially in edge devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。