arXiv:2411.02697cs.CVphysics.optics2024-11被引 18

光子编码器三通道并行卷积,算力降2万倍仍保高精度。

Transferable polychromatic optical encoder for neural networks

  • 用光学方式在成像时同步处理红绿蓝三通道卷积
  • 计算量减少约24,000倍,图像分类准确率达73.2%
  • 模型可直接迁移至ImageNet子集,无需重训练

人工神经网络(ANN)彻底改变了计算机视觉领域,实现了前所未有的性能。然而,用于图像处理的这些神经网络需要大量计算资源,常常阻碍实时运行。本文展示了一种可转移的多色光子编码器,能在图像捕获过程中同时对三个颜色通道执行卷积,相当于实现神经网络的若干初始卷积层。这种光学编码使计算操作减少约24,000倍,在自由空间光学系统中达到73.2%的先进分类准确率。此外,该模拟光学编码器在CIFAR-10数据上训练后,无需任何修改即可迁移到ImageNet子集High-10,仍保持适度准确率。结果表明,混合光/数字视觉系统具有潜力:光学前端可预处理环境场景,显著降低整个视觉系统的能耗和延迟。

原文摘要 · Abstract (English)

Artificial neural networks (ANNs) have fundamentally transformed the field of computer vision, providing unprecedented performance. However, these ANNs for image processing demand substantial computational resources, often hindering real-time operation. In this paper, we demonstrate an optical encoder that can perform convolution simultaneously in three color channels during the image capture, effectively implementing several initial convolutional layers of a ANN. Such an optical encoding results in ~24,000 times reduction in computational operations, with a state-of-the art classification accuracy (~73.2%) in free-space optical system. In addition, our analog optical encoder, trained for CIFAR-10 data, can be transferred to the ImageNet subset, High-10, without any modifications, and still exhibits moderate accuracy. Our results evidence the potential of hybrid optical/digital computer vision system in which the optical frontend can pre-process an ambient scene to reduce the energy and latency of the whole computer vision system.

光学计算神经网络加速跨数据集迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。