通过时间步混合提升静态到事件数据的知识迁移效率
Breaking the Modality Wall: Time-step Mixup for Efficient Spiking Knowledge Transfer from Static to Event Domain
- 设计时间步混合策略,利用脉冲神经网络异步特性融合图像与事件数据
- 在多个基准上实现优于现有方法的分类性能,最高提升4.2%准确率
- 适合研究低功耗视觉智能、跨模态知识迁移的工程师与学者
将事件相机与脉冲神经网络(SNN)结合可实现节能型视觉智能,但事件数据稀缺及DVS输出稀疏性限制了有效训练。以往从RGB到DVS的知识迁移因模态间分布差异大而表现不佳。本文提出时间步混合知识迁移(TMKT),一种基于概率时间步混合(TSM)的跨模态训练框架。TSM通过在不同时间步对齐RGB与DVS输入,生成序列内平滑的教学课程,降低梯度方差并稳定优化,理论分析支持其有效性。为利用TSM提供的辅助监督,TMKT引入两个轻量级模态感知目标:帧级源监督的模态感知引导(MAG)和序列级混合比例估计的混合比例感知(MRP),显式对齐时序特征与混合调度。该方法促进更平滑的知识迁移,缓解训练中的模态不匹配问题,在多个基准与多种SNN骨干网络上均取得优异表现。大量实验与消融分析验证了方法的有效性。
原文摘要 · Abstract (English)
The integration of event cameras and spiking neural networks (SNNs) promises energy-efficient visual intelligence, yet scarce event data and the sparsity of DVS outputs hinder effective training. Prior knowledge transfers from RGB to DVS often underperform because the distribution gap between modalities is substantial. In this work, we present Time-step Mixup Knowledge Transfer (TMKT), a cross-modal training framework with a probabilistic Time-step Mixup (TSM) strategy. TSM exploits the asynchronous nature of SNNs by interpolating RGB and DVS inputs at various time steps to produce a smooth curriculum within each sequence, which reduces gradient variance and stabilizes optimization with theoretical analysis. To employ auxiliary supervision from TSM, TMKT introduces two lightweight modality-aware objectives, Modality Aware Guidance (MAG) for per-frame source supervision and Mixup Ratio Perception (MRP) for sequence-level mix ratio estimation, which explicitly align temporal features with the mixing schedule. TMKT enables smoother knowledge transfer, helps mitigate modality mismatch during training, and achieves superior performance in spiking image classification tasks. Extensive experiments across diverse benchmarks and multiple SNN backbones, together with ablations, demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。