提升脉冲神经网络视频识别能力,通过优化运动频段实现高效动作识别
Fire on Motion: Optimizing Video Pass-bands for Efficient Spiking Action Recognition
- 设计PBO模块,动态优化脉冲网络的时序通带以增强运动信息捕捉
- 在UCF101上提升超10个百分点,在多模态与弱监督任务中表现显著
- 仅增加两个可学习参数,无需改变网络结构,适合部署于低功耗设备
脉冲神经网络(SNN)因能效高、生物合理性及天然时序处理能力,在视觉领域受到关注。然而,尽管具备时序处理潜力,现有研究仍集中于静态图像基准,导致SNN在动态视频任务上仍落后于人工神经网络(ANN)。本文诊断出根本原因:标准脉冲动力学表现为时序低通滤波,强调静态内容而抑制运动相关频段——动态任务中的关键信息集中在该频段。为解决此问题,提出可即插即用的通带优化器(PBO),仅引入两个可学习参数和轻量一致性约束,不改变网络结构,计算开销极小。PBO主动抑制对判别贡献小的静态成分,有效实现高通滤波,使脉冲活动聚焦于运动相关区域。在UCF101上取得超过10个百分点的性能提升;在更复杂的多模态动作识别与弱监督视频异常检测任务中也持续获得显著增益,为基于SNN的视频处理提供了新思路。
原文摘要 · Abstract (English)
Spiking neural networks (SNNs) have gained traction in vision due to their energy efficiency, bio-plausibility, and inherent temporal processing. Yet, despite this temporal capacity, most progress concentrates on static image benchmarks, and SNNs still underperform on dynamic video tasks compared to artificial neural networks (ANNs). In this work, we diagnose a fundamental pass-band mismatch: Standard spiking dynamics behave as a temporal low pass that emphasizes static content while attenuating motion bearing bands, where task relevant information concentrates in dynamic tasks. This phenomenon explains why SNNs can approach ANNs on static tasks yet fall behind on tasks that demand richer temporal understanding.To remedy this, we propose the Pass-Bands Optimizer (PBO), a plug-and-play module that optimizes the temporal pass-band toward task-relevant motion bands. PBO introduces only two learnable parameters, and a lightweight consistency constraint that preserves semantics and boundaries, incurring negligible computational overhead and requires no architectural changes. PBO deliberately suppresses static components that contribute little to discrimination, effectively high passing the stream so that spiking activity concentrates on motion bearing content. On UCF101, PBO yields over ten percentage points improvement. On more complex multi-modal action recognition and weakly supervised video anomaly detection, PBO delivers consistent and significant gains, offering a new perspective for SNN based video processing and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。