用事件相机原始点数据做动作识别,提升精度与鲁棒性
Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
- 通过掩码自编码器直接重建原始事件点,保留时空结构
- 新补丁生成算法提升事件点质量,减少噪声影响
- 首个在事件相机原始点上预训练的框架,适合低延迟场景
动态视觉传感器(DVS)是仿生设备,以异步事件形式捕捉视觉信息,具有高时间分辨率和低延迟。这些事件蕴含丰富的运动线索,可用于动作识别等计算机视觉任务。然而,现有基于DVS的动作识别方法在数据转换过程中会丢失时间信息,或受传感器缺陷及环境因素影响产生噪声和异常值。为此,我们提出一种新框架,旨在保留并利用事件数据的时空结构进行动作识别。该框架包含两个核心组件:1)点级事件掩码自编码器(MAE),通过重建被遮蔽的原始事件点数据,学习紧凑且有判别性的事件块表示;2)改进的事件点块生成算法,结合事件数据内点模型与点级数据增强技术,提升事件点块的质量与多样性。据我们所知,本方法首次将预训练引入事件相机原始点数据,并提出新型事件点块嵌入,使基于Transformer的模型能有效应用于事件相机。
原文摘要 · Abstract (English)
Dynamic vision sensors (DVS) are bio-inspired devices that capture visual information in the form of asynchronous events, which encode changes in pixel intensity with high temporal resolution and low latency. These events provide rich motion cues that can be exploited for various computer vision tasks, such as action recognition. However, most existing DVS-based action recognition methods lose temporal information during data transformation or suffer from noise and outliers caused by sensor imperfections or environmental factors. To address these challenges, we propose a novel framework that preserves and exploits the spatiotemporal structure of event data for action recognition. Our framework consists of two main components: 1) a point-wise event masked autoencoder (MAE) that learns a compact and discriminative representation of event patches by reconstructing them from masked raw event camera points data; 2) an improved event points patch generation algorithm that leverages an event data inlier model and point-wise data augmentation techniques to enhance the quality and diversity of event points patches. To the best of our knowledge, our approach introduces the pre-train method into event camera raw points data for the first time, and we propose a novel event points patch embedding to utilize transformer-based models on event cameras.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。