arXiv:2505.20834cs.CVcs.NE2025-05NeurIPS被引 6

首个全脉冲网络实现帧与事件流统一追踪,兼顾精度与能效。

Fully Spiking Neural Networks for Unified Frame-Event Object Tracking

  • 全脉冲架构融合卷积局部特征与Transformer全局建模。
  • 在多个基准上精度超越现有方法,功耗显著降低。
  • 适合低功耗实时追踪场景,如嵌入式视觉系统。

图像与事件流的融合为复杂环境下的鲁棒视觉目标追踪提供了新路径。然而,现有融合方法虽性能优异,却带来巨大计算开销,且难以高效提取事件流中稀疏、异步的信息,未能发挥脉冲驱动范式的节能优势。为此,我们提出首个全脉冲帧-事件追踪框架SpikeFET,实现了在脉冲范式下卷积局部特征提取与Transformer全局建模的协同融合。为克服卷积填充导致的平移不变性退化,引入随机拼贴模块(RPM),通过随机空间重组与可学习类型编码消除位置偏差,同时保留残差结构。此外,提出时空正则化(STR)策略,通过强制时序模板特征在潜在空间中的时空一致性,缓解因特征不对称导致的相似性度量退化。大量实验表明,该框架在多个基准上达到更优追踪精度,同时显著降低功耗,实现了性能与效率的最优平衡。

原文摘要 · Abstract (English)

The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, asynchronous information from event streams, failing to leverage the energy-efficient advantages of event-driven spiking paradigms. To address this challenge, we propose the first fully Spiking Frame-Event Tracking framework called SpikeFET. This network achieves synergistic integration of convolutional local feature extraction and Transformer-based global modeling within the spiking paradigm, effectively fusing frame and event data. To overcome the degradation of translation invariance caused by convolutional padding, we introduce a Random Patchwork Module (RPM) that eliminates positional bias through randomized spatial reorganization and learnable type encoding while preserving residual structures. Furthermore, we propose a Spatial-Temporal Regularization (STR) strategy that overcomes similarity metric degradation from asymmetric features by enforcing spatio-temporal consistency among temporal template features in latent space. Extensive experiments across multiple benchmarks demonstrate that the proposed framework achieves superior tracking accuracy over existing methods while significantly reducing power consumption, attaining an optimal balance between performance and efficiency.

脉冲神经网络目标追踪事件相机低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。