arXiv:2410.12416cs.SDcs.AI2024-10被引 6

用分段平均池化提升语音情感识别,更专注有效语段。

Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features

  • 提出SAP方法,只保留有情感信息的语音片段
  • 在IEMOCAP和KEMDy19上达到最优准确率
  • 适合做语音情感分析的研究者参考

语音情感识别(SER)旨在分析语音中表达的人类情绪。自监督学习(SSL)通过大量无标签音频数据学习有意义的表示,为SER提供了新路径。然而,现有方法依赖全局平均池化(GAP)提取特征,对语音与非语音段同等处理,导致有用信息被无关内容稀释。为此,本文提出分段平均池化(SAP),仅聚焦于具有信息量的语音片段,忽略非语音部分。同时结合GAP与SAP,兼顾整体信号与关键片段信息,显著提升性能。实验表明,在英语IEMOCAP数据集上达到当前最优结果,在韩语KEMDy19数据集上无论加权还是未加权准确率均领先。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) analyzes human emotions expressed through speech. Self-supervised learning (SSL) offers a promising approach to SER by learning meaningful representations from a large amount of unlabeled audio data. However, existing SSL-based methods rely on Global Average Pooling (GAP) to represent audio signals, treating speech and non-speech segments equally. This can lead to dilution of informative speech features by irrelevant non-speech information. To address this, the paper proposes Segmental Average Pooling (SAP), which selectively focuses on informative speech segments while ignoring non-speech segments. By applying both GAP and SAP to SSL features, our approach utilizes overall speech signal information from GAP and specific information from SAP, leading to improved SER performance. Experiments show state-of-the-art results on the IEMOCAP for English and superior performance on KEMDy19 for Korean datasets in both unweighted and weighted accuracies.

语音情感识别自监督学习特征池化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。