arXiv:2505.19938cs.CV2025-05中稿 · IEEE TCSVT被引 8

用脉冲神经网络分离语义与动态运动,提升跨模态零样本视频识别效果

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

  • 双流架构分离语义与稀疏运动信息,通过事件图像增强运动捕捉
  • 引入差异分析块建模音频运动,动态调整神经元阈值提升鲁棒性
  • 在主流数据集上零样本识别准确率提升39.9%,适合多模态动作识别研究

音频-视觉零样本学习(ZSL)旨在对训练中未见类别的视频进行分类。现有方法常受背景场景偏差和运动细节不足的影响。本文提出一种新型双流多时尺度运动解耦脉冲变压器(MDST++),将上下文语义信息与稀疏动态运动信息解耦。通过递归联合学习单元提取语义信息并捕获多模态间联合知识以理解动作环境。将RGB图像转换为事件数据,更精确地捕捉运动信息并缓解背景偏差。此外,引入差异分析块建模音频运动信息。为增强脉冲神经网络(SNN)在提取时空与运动线索上的鲁棒性,基于全局运动和语义信息动态调整漏电积分-放电神经元阈值。实验验证了MDST++的有效性,在主流基准测试中持续优于当前最优方法。结合运动与多时尺度信息,显著提升HM与ZSL准确率26.2%和39.9%。

原文摘要 · Abstract (English)

Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scene biases and inadequate motion detail. This paper proposes a novel dual-stream Multi-Timescale Motion-Decoupled Spiking Transformer (MDST++), which decouples contextual semantic information and sparse dynamic motion information. The recurrent joint learning unit is proposed to extract contextual semantic information and capture joint knowledge across various modalities to understand the environment of actions. By converting RGB images to events, our method captures motion information more accurately and mitigates background scene biases. Moreover, we introduce a discrepancy analysis block to model audio motion information. To enhance the robustness of SNNs in extracting temporal and motion cues, we dynamically adjust the threshold of Leaky Integrate-and-Fire neurons based on global motion and contextual semantic information. Our experiments validate the effectiveness of MDST++, demonstrating their consistent superiority over state-of-the-art methods on mainstream benchmarks. Additionally, incorporating motion and multi-timescale information significantly improves HM and ZSL accuracy by 26.2\% and 39.9\%.

零样本学习脉冲神经网络多模态动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。