arXiv:2601.17391cs.CV2026-01被引 1

提出高效事件相机动作识别新框架,精度提升且参数减少三成。

SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition

  • 通过平移不变的密集事件转换构建时空多视角表征
  • 动态双分支融合使不同视角运动特征互补,精度提升超10%
  • 仿生时间扭曲增强模拟真实动作速度变化,适合低功耗场景

事件相机动作识别(EAR)具备隐私保护和高效的优势,其中时间运动动态至关重要。现有时空多视角表征学习(SMVRL)方法虽通过沿空间轴H和W投影事件展现潜力,但受限于平移敏感的空间分箱表示和简单的早期拼接融合结构。本文重新审视EAR中SMVRL的关键设计环节,提出:(i) 基于平移不变的稀疏事件密集转换的合理时空多视角表征;(ii) 双分支动态融合架构,建模不同视角间样本级运动特征的互补性;(iii) 生物启发的时间扭曲增强,模拟真实人类动作的速度变异性。在HARDVS、DailyDVS-200和THU-EACT-50-CHL三个挑战性EAR数据集上,相比现有SMVRL EOR方法,分别实现+7.0%、+10.7%和+10.2%的Top-1准确率提升,同时参数减少30.1%,计算量降低35.7%,确立了该框架作为新型高效EAR范式。

原文摘要 · Abstract (English)

Event cameras action recognition (EAR) offers compelling privacy-protecting and efficiency advantages, where temporal motion dynamics is of great importance. Existing spatiotemporal multi-view representation learning (SMVRL) methods for event-based object recognition (EOR) offer promising solutions by projecting H-W-T events along spatial axis H and W, yet are limited by its translation-variant spatial binning representation and naive early concatenation fusion architecture. This paper reexamines the key SMVRL design stages for EAR and propose: (i) a principled spatiotemporal multi-view representation through translation-invariant dense conversion of sparse events, (ii) a dual-branch, dynamic fusion architecture that models sample-wise complementarity between motion features from different views, and (iii) a bio-inspired temporal warping augmentation that mimics speed variability of real-world human actions. On three challenging EAR datasets of HARDVS, DailyDVS-200 and THU-EACT-50-CHL, we show +7.0%, +10.7%, and +10.2% Top-1 accuracy gains over existing SMVRL EOR method with surprising 30.1% reduced parameters and 35.7% lower computations, establishing our framework as a novel and powerful EAR paradigm.

事件相机动作识别多视角学习高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。