arXiv:2603.16338cs.CV2026-03

用自监督学习让脉冲神经网络从无标签事件数据中学出强视觉表征。

SpikeCLR: Contrastive Self-Supervised Learning for Few-Shot Event-Based Vision using Spiking Neural Networks

  • 设计脉冲神经网络的对比自监督框架,适配事件数据特性。
  • 少样本场景下性能超有监督训练,跨数据集迁移有效。
  • 融合时空极性增强,提升对事件流时序不变性的建模能力。

事件视觉传感器在高速感知中具有微秒级时间分辨率、高动态范围和低功耗的优势。结合脉冲神经网络(SNNs)可在类脑硬件上实现能效优化的嵌入式应用。然而,其潜力受限于大规模标注数据的缺乏。本文提出SpikeCLR,一种面向事件数据的对比自监督学习框架,使SNN能够从未标记事件流中学习鲁棒视觉表示。通过代理梯度训练将图像域方法迁移至脉冲域,并引入空间、时间与极性变换等事件特异性增强策略。在CIFAR10-DVS、N-Caltech101、N-MNIST和DVS-Gesture多个基准上验证,自监督预训练+微调在少样本与半监督设置中均优于有监督学习。消融实验表明,时空增强的联合使用对学习有效的时空不变性至关重要。此外,学习到的表征具备跨数据集迁移能力,推动了标签稀缺环境下高性能事件视觉模型的发展。

原文摘要 · Abstract (English)

Event-based vision sensors provide significant advantages for high-speed perception, including microsecond temporal resolution, high dynamic range, and low power consumption. When combined with Spiking Neural Networks (SNNs), they can be deployed on neuromorphic hardware, enabling energy-efficient applications on embedded systems. However, this potential is severely limited by the scarcity of large-scale labeled datasets required to effectively train such models. In this work, we introduce SpikeCLR, a contrastive self-supervised learning framework that enables SNNs to learn robust visual representations from unlabeled event data. We adapt prior frame-based methods to the spiking domain using surrogate gradient training and introduce a suite of event-specific augmentations that leverage spatial, temporal, and polarity transformations. Through extensive experiments on CIFAR10-DVS, N-Caltech101, N-MNIST, and DVS-Gesture benchmarks, we demonstrate that self-supervised pretraining with subsequent fine-tuning outperforms supervised learning in low-data regimes, achieving consistent gains in few-shot and semi-supervised settings. Our ablation studies reveal that combining spatial and temporal augmentations is critical for learning effective spatio-temporal invariances in event data. We further show that learned representations transfer across datasets, contributing to efforts for powerful event-based models in label-scarce settings.

事件视觉脉冲神经网络自监督学习少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。