arXiv:2602.08240cs.AIcs.SD2026-02被引 1

用提示调优让脉冲神经网络高效识别语音情绪,适合边缘设备部署。

PTS-SNN: A Prompt-Tuned Temporal Shift Spiking Neural Networks for Efficient Speech Emotion Recognition

  • 通过时序移位编码器和上下文感知电压校准,适配预训练语音模型与脉冲神经网络。
  • 在IEMOCAP数据集上达73.34%准确率,仅需119万参数和0.35毫焦能量。
  • 参数高效、低功耗,适合资源受限的语音情绪识别场景。

语音情绪识别(SER)广泛应用于人机交互,但传统模型计算成本高,难以在资源受限的边缘设备上部署。脉冲神经网络(SNN)因其事件驱动特性具有能耗优势,但其与连续自监督学习(SSL)表征的融合受分布不匹配制约——高动态范围嵌入会降低阈值型神经元的信息编码能力。为此,本文提出提示调优脉冲神经网络(PTS-SNN),一种参数高效的类脑适配框架,实现冻结的SSL主干与脉冲动力学的对齐。具体地,引入无参数通道时序移位编码器,捕捉局部时序依赖,建立稳定特征基础;设计上下文感知膜电位校准策略,利用脉冲稀疏线性注意力模块聚合全局语义上下文生成可学习软提示,动态调节可变泄漏积分-放电(PLIF)神经元的偏置电压,有效将异构输入分布聚焦于响应区间,缓解功能静默或饱和问题。在五个多语言数据集(如IEMOCAP、CASIA、EMODB)上的实验表明,PTS-SNN在IEMOCAP上达到73.34%准确率,媲美先进人工神经网络(ANNs),同时仅需119万可训练参数与每样本0.35毫焦推理能量。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) is widely deployed in Human-Computer Interaction, yet the high computational cost of conventional models hinders their implementation on resource-constrained edge devices. Spiking Neural Networks (SNNs) offer an energy-efficient alternative due to their event-driven nature; however, their integration with continuous Self-Supervised Learning (SSL) representations is fundamentally challenged by distribution mismatch, where high-dynamic-range embeddings degrade the information coding capacity of threshold-based neurons. To resolve this, we propose Prompt-Tuned Spiking Neural Networks (PTS-SNN), a parameter-efficient neuromorphic adaptation framework that aligns frozen SSL backbones with spiking dynamics. Specifically, we introduce a Temporal Shift Spiking Encoder to capture local temporal dependencies via parameter-free channel shifts, establishing a stable feature basis. To bridge the domain gap, we devise a Context-Aware Membrane Potential Calibration strategy. This mechanism leverages a Spiking Sparse Linear Attention module to aggregate global semantic context into learnable soft prompts, which dynamically regulate the bias voltages of Parametric Leaky Integrate-and-Fire (PLIF) neurons. This regulation effectively centers the heterogeneous input distribution within the responsive firing range, mitigating functional silence or saturation. Extensive experiments on five multilingual datasets (e.g., IEMOCAP, CASIA, EMODB) demonstrate that PTS-SNN achieves 73.34\% accuracy on IEMOCAP, comparable to competitive Artificial Neural Networks (ANNs), while requiring only 1.19M trainable parameters and 0.35 mJ inference energy per sample.

脉冲神经网络语音情绪识别边缘计算低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。