用语言模型思路提升事件相机的异步特征表达能力
Maximizing Asynchronicity in Event-based Neural Networks
- 借鉴语言建模思想,用线性注意力和自监督学习构建异步特征
- 在手势与车辆数据集上超越已有方法,首次在Gen1检测任务达0.477mAP
- 适合需要低延迟实时视觉的应用场景
事件相机以高时间分辨率、低延迟和极少冗余提供视觉数据,但其异步稀疏序列特性难以适配传统张量型机器学习。现有异步转同步(A2S)方法虽尝试弥合差距,却常牺牲表达力与泛化性。本文提出EVA(EVent Asynchronous feature learning),一种新型A2S框架,能生成高度表达且通用的逐事件特征。受事件与语言类比启发,EVA创新性地引入线性注意力与自监督学习技术。实验表明,EVA在识别任务(DVS128-Gesture和N-Cars)上优于先前A2S方法,并首次实现对复杂检测任务的成功处理,在Gen1数据集上取得0.477 mAP的性能。结果验证了EVA在推动实时事件视觉应用方面的潜力。
原文摘要 · Abstract (English)
Event cameras deliver visual data with high temporal resolution, low latency, and minimal redundancy, yet their asynchronous, sparse sequential nature challenges standard tensor-based machine learning (ML). While the recent asynchronous-to-synchronous (A2S) paradigm aims to bridge this gap by asynchronously encoding events into learned features for ML pipelines, existing A2S approaches often sacrifice expressivity and generalizability compared to dense, synchronous methods. This paper introduces EVA (EVent Asynchronous feature learning), a novel A2S framework to generate highly expressive and generalizable event-by-event features. Inspired by the analogy between events and language, EVA uniquely adapts advances from language modeling in linear attention and self-supervised learning for its construction. In demonstration, EVA outperforms prior A2S methods on recognition tasks (DVS128-Gesture and N-Cars), and represents the first A2S framework to successfully master demanding detection tasks, achieving a 0.477 mAP on the Gen1 dataset. These results underscore EVA's potential for advancing real-time event-based vision applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。