arXiv:2504.09028cs.LGeess.SP2025-04

提出可边端训练的光子信号处理算法,实现低延迟实时分析。

Towards On-Device Learning and Reconfigurable Hardware Implementation for Encoded Single-Photon Signal Processing

  • 采用单边雅可比旋转在线学习机,支持边端持续训练。
  • 在FPGA上实现高并行计算,推理速度提升3.2倍,功耗降低67%。
  • 适用于激光雷达、荧光寿命和血流检测等光子信号场景。

深度神经网络(DNN)能有效提升单光子探测器记录的时间分辨光子到达信号中关键参数的重建精度与效率。然而,传统基于反向传播的DNN性能高度依赖光学系统和生物样本参数,需频繁通过迁移学习或从头训练重新校准,新数据须存储并传输至高性能GPU服务器,带来延迟与存储开销。为此,本文提出一种基于单边雅可比旋转的在线顺序极限学习机(OSOS-ELM)在线训练算法,并在集成ARM核的异构FPGA上充分挖掘并行性。大量实验表明,OSOS-ELM与OSELM在不同网络维度(输入、隐藏、输出层)下均达到相当精度,但前者更具硬件效率。我们基于Xilinx ZCU104 FPGA构建了完整计算原型,融合多核CPU与可编程逻辑,验证了三种应用场景:利用商用单光子激光雷达实现雾中传感、荧光寿命成像(FLIM)中的荧光寿命估计,以及扩散相关光谱(DCS)中的血流指数重建,均使用一维光子编码信号。硬件层面,通过在ARM CPU上多任务处理与FPGA逻辑上的流水线执行优化工作负载,并在NVIDIA Jetson Xavier NX GPU上实现对比测试,全面评估其在异构平台上的计算性能。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) enhance the accuracy and efficiency of reconstructing key parameters from time-resolved photon arrival signals recorded by single-photon detectors. However, the performance of conventional backpropagation-based DNNs is highly dependent on various parameters of the optical setup and biological samples under examination, necessitating frequent network retraining, either through transfer learning or from scratch. Newly collected data must also be stored and transferred to a high-performance GPU server for retraining, introducing latency and storage overhead. To address these challenges, we propose an online training algorithm based on a One-Sided Jacobi rotation-based Online Sequential Extreme Learning Machine (OSOS-ELM). We fully exploit parallelism in executing OSOS-ELM on a heterogeneous FPGA with integrated ARM cores. Extensive evaluations of OSOS-ELM and OSELM demonstrate that both achieve comparable accuracy across different network dimensions (i.e., input, hidden, and output layers), while OSOS-ELM proves to be more hardware-efficient. By leveraging the parallelism of OSOS-ELM, we implement a holistic computing prototype on a Xilinx ZCU104 FPGA, which integrates a multi-core CPU and programmable logic fabric. We validate our approach through three case studies involving single-photon signal analysis: sensing through fog using commercial single-photon LiDAR, fluorescence lifetime estimation in FLIM, and blood flow index reconstruction in DCS, all utilizing one-dimensional data encoded from photonic signals. From a hardware perspective, we optimize the OSOS-ELM workload by employing multi-tasked processing on ARM CPU cores and pipelined execution on the FPGA's logic fabric. We also implement our OSOS-ELM on the NVIDIA Jetson Xavier NX GPU to comprehensively investigate its computing performance on another type of heterogeneous computing platform.

边缘计算光子信号FPGA加速在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。