arXiv:2510.26614cs.CVcs.RO2025-10被引 2

提出新型事件令牌化方法,保留事件相机的异步稀疏特性。

Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras

  • 用脉冲补丁实现事件的异步稀疏令牌化,保持原始数据特性。
  • 推理速度比体素快3.4倍,比帧快10.4倍,准确率不降反升。
  • 适合低延迟视觉任务,如手势识别与目标检测应用。

我们提出事件的令牌化方法,并设计了专为事件相机优化的分块脉冲(Spiking Patches)分词器。针对异步且空间稀疏的事件流,目标是获得保留这些特性的事件表示。现有方法将事件表示为帧或体素,虽精度高,但均为同步且破坏空间稀疏性。脉冲补丁可保留事件相机的独特性质,在实验中证明其性能不降反升。我们在手势识别和目标检测任务上,使用GNN、PCN和Transformer评估该分词器。相比体素令牌,脉冲补丁推理速度提升最高达3.4倍;相比帧,最快达10.4倍。在保持相同甚至更高准确率的同时,手势识别绝对提升达3.8,目标检测达1.4。因此,令牌化为事件视觉开辟新方向,推动真正保留事件相机特性的方法发展。

原文摘要 · Abstract (English)

We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous and spatially sparse events, our goal is to discover an event representation that preserves these properties. Prior works have represented events as frames or as voxels. However, while these representations yield high accuracy, both frames and voxels are synchronous and decrease the spatial sparsity. Spiking Patches gives the means to preserve the unique properties of event cameras and we show in our experiments that this comes without sacrificing accuracy. We evaluate our tokenizer using a GNN, PCN, and a Transformer on gesture recognition and object detection. Tokens from Spiking Patches yield inference times that are up to 3.4x faster than voxel-based tokens and up to 10.4x faster than frames. We achieve this while matching their accuracy and even surpassing in some cases with absolute improvements up to 3.8 for gesture recognition and up to 1.4 for object detection. Thus, tokenization constitutes a novel direction in event-based vision and marks a step towards methods that preserve the properties of event cameras.

事件相机稀疏表示高效推理脉冲神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。