arXiv:2411.18328cs.CV2024-11被引 5

融合帧与点的协同机制,提升事件相机动作识别精度与效率

EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond

  • 设计帧点协同框架,兼顾密集帧与稀疏点的特性
  • 在SeAct和HARDVS上分别提升5.17%和7.01%
  • 适合事件相机、低功耗视觉任务的研究者

事件相机动作识别(EAR)相比传统方法具有高时间分辨率和隐私保护优势。现有主流方法分为两类:将原始事件流投影为稠密事件帧并使用强帧模型,或直接用轻量点模型处理稀疏事件点。但二者均忽视了异步事件数据特有的稠密时序与稀疏空间特性。本文提出协同感知框架EventCrab,巧妙结合‘轻’帧模型与‘重’点模型,在准确率与效率间取得平衡。进一步构建帧-文本-点联合表征空间,以弥合不同表示差异。针对事件点的时空关系,设计两种增强策略:①脉冲式上下文学习器(SCL),从原始事件流中提取上下文信息;②事件点编码器(EPE),采用希尔伯特扫描方式挖掘长时序空间特征。在四个数据集上的实验表明,EventCrab表现显著,尤其在SeAct上提升5.17%,HARDVS上提升7.01%。

原文摘要 · Abstract (English)

Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current leading EAR solutions typically follow two regimes: project unconstructed event streams into dense constructed event frames and adopt powerful frame-specific networks, or employ lightweight point-specific networks to handle sparse unconstructed event points directly. However, such two regimes are blind to a fundamental issue: failing to accommodate the unique dense temporal and sparse spatial properties of asynchronous event data. In this article, we present a synergy-aware framework, i.e., EventCrab, that adeptly integrates the "lighter" frame-specific networks for dense event frames with the "heavier" point-specific networks for sparse event points, balancing accuracy and efficiency. Furthermore, we establish a joint frame-text-point representation space to bridge distinct event frames and points. In specific, to better exploit the unique spatiotemporal relationships inherent in asynchronous event points, we devise two strategies for the "heavier" point-specific embedding: i) a Spiking-like Context Learner (SCL) that extracts contextualized event points from raw event streams. ii) an Event Point Encoder (EPE) that further explores event-point long spatiotemporal features in a Hilbert-scan way. Experiments on four datasets demonstrate the significant performance of our proposed EventCrab, particularly gaining improvements of 5.17% on SeAct and 7.01% on HARDVS.

事件相机动作识别协同建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。