arXiv:2412.16904cs.SDeess.AS2024-12中稿 · ICASSP 2025被引 24

提出时空双域协同的语音情感识别框架,兼顾效率与表现

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

  • 设计时频双域感知模块,同步捕捉语音时序与频域特征
  • 在IEMOCAP和MELD数据集上达到更高准确率,模型更小延迟更低
  • 适合追求实时性与高精度的语音交互系统开发者使用

语音情感识别(SER)在人机交互中对提升用户体验至关重要。然而,现有方法过度依赖时域分析,忽视了频域包络结构中同样重要的情感线索。为此,我们提出TF-Mamba,一种新型多域框架,同时捕捉时域与频域的情感表达。具体地,设计时频马比块以提取兼具时序与频域感知的特征,实现计算效率与模型表达力的最佳平衡;此外,引入复数度量-距离三元组损失(CMDT),使模型更精准捕捉代表性情感线索。在IEMOCAP和MELD数据集上的大量实验表明,TF-Mamba在模型尺寸与推理延迟方面优于现有方法,为未来SER应用提供了更实用的解决方案。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emotion recognition. To overcome this limitation, we propose TF-Mamba, a novel multi-domain framework that captures emotional expressions in both temporal and frequency dimensions.Concretely, we propose a temporal-frequency mamba block to extract temporal- and frequency-aware emotional features, achieving an optimal balance between computational efficiency and model expressiveness. Besides, we design a Complex Metric-Distance Triplet (CMDT) loss to enable the model to capture representative emotional clues for SER. Extensive experiments on the IEMOCAP and MELD datasets show that TF-Mamba surpasses existing methods in terms of model size and latency, providing a more practical solution for future SER applications.

语音情感识别时频融合高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。