用全息投影+频域门控,让事件相机动作识别更快更准
HoloEv-Net: Efficient Event-based Action Recognition via Holographic Spatial Embedding and Global Spectral Gating
- 用2D全息表示替代密集体素,压缩空间冗余
- 频域全局门控增强运动模式捕捉,参数几乎不增加
- 轻量版效率提升300倍,适合边缘设备部署
事件相机因高时间分辨率和高动态范围,使事件动作识别备受关注。但现有方法普遍存在(i)密集体素表示的计算冗余,(ii)多分支结构的结构冗余,以及(iii)对频谱信息利用不足的问题。为此,我们提出高效事件动作识别框架HoloEv-Net。首先,为同时解决表示与结构冗余,引入紧凑全息时空表示(CHSR)。不同于高耗能体素网格,CHSR将水平空间信息隐式嵌入时间-高度(T-H)视图,在2D表示中有效保留3D时空上下文。其次,为挖掘被忽视的频谱线索,设计全局频谱门控(GSG)模块。通过快速傅里叶变换(FFT)在频域实现全局令牌混合,显著增强表征能力且参数开销极低。大量实验表明框架可扩展性强且有效:HoloEv-Net-Base在THU-EACT-50-CHL、HARDVS和DailyDVS-200上分别优于现有方法10.29%、1.71%和6.25%;其轻量版HoloEv-Net-Small在保持高精度的同时,参数减少5.4倍,计算量降低300倍,延迟下降2.4倍,具备边缘部署潜力。
原文摘要 · Abstract (English)
Event-based Action Recognition (EAR) has attracted significant attention due to the high temporal resolution and high dynamic range of event cameras. However, existing methods typically suffer from (i) the computational redundancy of dense voxel representations, (ii) structural redundancy inherent in multi-branch architectures, and (iii) the under-utilization of spectral information in capturing global motion patterns. To address these challenges, we propose an efficient EAR framework named HoloEv-Net. First, to simultaneously tackle representation and structural redundancies, we introduce a Compact Holographic Spatiotemporal Representation (CHSR). Departing from computationally expensive voxel grids, CHSR implicitly embeds horizontal spatial cues into the Time-Height (T-H) view, effectively preserving 3D spatiotemporal contexts within a 2D representation. Second, to exploit the neglected spectral cues, we design a Global Spectral Gating (GSG) module. By leveraging the Fast Fourier Transform (FFT) for global token mixing in the frequency domain, GSG enhances the representation capability with negligible parameter overhead. Extensive experiments demonstrate the scalability and effectiveness of our framework. Specifically, HoloEv-Net-Base achieves state-of-the-art performance on THU-EACT-50-CHL, HARDVS and DailyDVS-200, outperforming existing methods by 10.29%, 1.71% and 6.25%, respectively. Furthermore, our lightweight variant, HoloEv-Net-Small, delivers highly competitive accuracy while offering extreme efficiency, reducing parameters by 5.4 times, FLOPs by 300times, and latency by 2.4times compared to heavy baselines, demonstrating its potential for edge deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。