arXiv:2604.24800cs.AReess.IV2026-04

用原子光子系统加速3D卷积,实现超高速视频分类

Opto-Atomic Spatio-Temporal Holographic Correlators for High-Speed 3D CNNs

  • 将3D卷积转移至原子光子相关器,同步处理空间与时间信息
  • 在动作识别数据集上达59.72%准确率,支持30×40像素×8帧的大核并行
  • 可实现最高12.5万帧/秒的处理速度,适合实时视频分析场景

三维卷积神经网络(3D CNNs)在视频识别中表现优异,能同时处理空间与时间特征。但其计算复杂度呈立方增长,对传统硅基硬件造成显著时延与能耗挑战。为此,我们提出一种混合光电架构,将计算密集的3D卷积层交由光子-原子时空全息相关器(STHC)执行。该系统利用冷铷-85原子阵列存储时间信息为原子相干态,并结合传统二维空间相关器,实现时空同步相关。在四类人体动作数据集上的实验表明,使用30×40像素的空间核与8帧的时间跨度时,分类准确率达59.72%,且理论运行速度可达每秒12.5万帧。该方法为通过混合架构实现视频分类的大幅加速提供了新路径。

原文摘要 · Abstract (English)

Three-dimensional convolutional neural networks (3D CNNs) have demonstrated remarkable performance in video recognition tasks by processing both spatial and temporal features. However, the cubic scaling of computational complexity poses significant time and energy efficiency challenges for conventional silicon-based hardware. To address this, we propose a hybrid optoelectronic architecture that delegates the computationally intensive 3D convolutional layer to an opto-atomic Spatio-temporal Holographic Correlator (STHC). This system stores temporal information as atomic coherence in an array of inhomogeneously broadened cold Rubidium-85 atoms and combines a traditional 2D spatial correlator to perform correlation in both space and time simultaneously. Our results on a four-class human action dataset demonstrate a classification accuracy of 59.72% using parallel large-scale kernels (30X40 pixels spatially, 8 frames temporally), with potential operating speeds projected up to 125,000 frames per second. This approach offers a pathway to massively accelerated video classification through a hybrid architecture.

3D卷积光子计算原子系统视频分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。