用脉冲神经网络实现高精度眼动追踪,省电又实时。
GazeSCRNN: Event-based Near-eye Gaze Tracking using a Spiking Neural Network
- 用脉冲神经网络处理动态视觉传感器事件流,捕捉快速眼动。
- 在EV-Eye数据集上达到6.034°平均角度误差,2.094mm平均瞳孔误差。
- 首次证明脉冲网络可高效用于眼动追踪,适合低功耗设备应用。
本文提出GazeSCRNN,一种用于事件相机近眼眼动追踪的新型脉冲卷积循环神经网络。利用动态视觉传感器(DVS)的高时间分辨率、低功耗特性及与事件系统的兼容性,GazeSCRNN采用脉冲神经网络(SNN)克服传统眼动追踪系统在捕捉动态运动时的局限。模型通过自适应漏电整合放电(ALIF)神经元与优化的时空混合架构,处理来自DVS的事件流。在EV-Eye数据集上的大量实验表明,该模型能准确预测眼动向量。消融实验揭示了ALIF神经元、动态事件帧划分及前向传播时序训练等技术对性能提升的关键作用。最优模型实现6.034°的平均角度误差(MAE)和2.094mm的平均瞳孔误差(MPE)。本工作首次验证了脉冲神经网络在事件基眼动追踪中的可行性,为后续研究指明关键挑战与改进方向。
原文摘要 · Abstract (English)
This work introduces GazeSCRNN, a novel spiking convolutional recurrent neural network designed for event-based near-eye gaze tracking. Leveraging the high temporal resolution, energy efficiency, and compatibility of Dynamic Vision Sensor (DVS) cameras with event-based systems, GazeSCRNN uses a spiking neural network (SNN) to address the limitations of traditional gaze-tracking systems in capturing dynamic movements. The proposed model processes event streams from DVS cameras using Adaptive Leaky-Integrate-and-Fire (ALIF) neurons and a hybrid architecture optimized for spatio-temporal data. Extensive evaluations on the EV-Eye dataset demonstrate the model's accuracy in predicting gaze vectors. In addition, we conducted ablation studies to reveal the importance of the ALIF neurons, dynamic event framing, and training techniques, such as Forward-Propagation-Through-Time, in enhancing overall system performance. The most accurate model achieved a Mean Angle Error (MAE) of 6.034° and a Mean Pupil Error (MPE) of 2.094 mm. Consequently, this work is pioneering in demonstrating the feasibility of using SNNs for event-based gaze tracking, while shedding light on critical challenges and opportunities for further improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。