arXiv:2509.12959cs.CV2025-09

通过时间步混合提升视觉到事件域的知识迁移效率

Time-step Mixup for Efficient Spiking Knowledge Transfer from Appearance to Event Domain

  • 在不同时间步混合RGB与事件数据,利用脉冲网络异步特性
  • 引入模态感知辅助目标,实现跨模态标签混合与判别增强
  • 显著缓解模态差异,提升事件图像分类性能,适合低功耗视觉任务

事件相机与脉冲神经网络的结合为节能视觉处理带来巨大潜力。然而,事件数据有限且DVS输出稀疏,给有效训练带来挑战。尽管已有工作尝试将RGB数据集中的语义知识迁移到DVS,但常忽视两者间显著的分布差异。本文提出时间步混合知识迁移(TMKT),一种新颖的细粒度混合策略,通过在不同时间步对齐RGB与DVS输入,利用脉冲网络的异步特性。为进一步支持跨模态标签混合,我们引入模态感知辅助学习目标,辅助时间步混合过程并增强模型在不同模态间的判别能力。该方法实现更平滑的知识迁移,缓解训练中的模态偏移,在多个数据集上均取得优异的脉冲图像分类性能。实验验证了方法的有效性。代码将在双盲评审后公开。

原文摘要 · Abstract (English)

The integration of event cameras and spiking neural networks holds great promise for energy-efficient visual processing. However, the limited availability of event data and the sparse nature of DVS outputs pose challenges for effective training. Although some prior work has attempted to transfer semantic knowledge from RGB datasets to DVS, they often overlook the significant distribution gap between the two modalities. In this paper, we propose Time-step Mixup knowledge transfer (TMKT), a novel fine-grained mixing strategy that exploits the asynchronous nature of SNNs by interpolating RGB and DVS inputs at various time-steps. To enable label mixing in cross-modal scenarios, we further introduce modality-aware auxiliary learning objectives. These objectives support the time-step mixup process and enhance the model's ability to discriminate effectively across different modalities. Our approach enables smoother knowledge transfer, alleviates modality shift during training, and achieves superior performance in spiking image classification tasks. Extensive experiments demonstrate the effectiveness of our method across multiple datasets. The code will be released after the double-blind review process.

脉冲神经网络知识迁移事件相机跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。