arXiv:2410.00390eess.AS2024-10被引 25

提出多尺度时序变换器,提升语音情感识别精度并降低计算开销。

Multi-Scale Temporal Transformer For Speech Emotion Recognition

  • 设计多尺度时序特征算子,捕捉不同局部时间片段的情感特征。
  • 在IEMOCAP、MELD、CREMAD上超越现有模型,同时减少计算量。
  • 适合需要高效高精度语音情感分析的应用场景。

语音情感识别在人机交互系统中至关重要。尽管近期优化的Transformer已在该任务中取得成功,但现有架构更关注全局信息,计算开销大。事实上,输入语音中存在丰富的局部情感表征。为此,本文提出多尺度时序变换器(MSTR),包含三个核心组件:(1) 多尺度时序特征算子,(2) 分形自注意力模块,(3) 尺度混合模块。三者协同增强Transformer对多尺度局部情感表征的学习能力。实验表明,MSTR在IEMOCAP、MELD和CREMAD三个语音情感数据集上显著优于原始Transformer及其他先进方法,同时大幅降低计算成本。

原文摘要 · Abstract (English)

Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures focus more on global information and require large computation. On the other hand, abundant speech emotional representations exist locally on different parts of the input speech. To tackle these problems, we propose a Multi-Scale TRansfomer (MSTR) for speech emotion recognition. It comprises of three main components: (1) a multi-scale temporal feature operator, (2) a fractal self-attention module, and (3) a scale mixer module. These three components can effectively enhance the transformer's ability to learn multi-scale local emotion representations. Experimental results demonstrate that the proposed MSTR model significantly outperforms a vanilla Transformer and other state-of-the-art methods across three speech emotion datasets: IEMOCAP, MELD and, CREMAD. In addition, it can greatly reduce the computational cost.

语音情感识别Transformer多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。