arXiv:2409.04007cs.SDcs.AI2024-09被引 7

通过优化预处理与轻量注意力结构,提升小样本语音情感识别性能

Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition

  • 设计6层CNN+高效通道注意力,以少参数提升特征表达
  • 高频时频分辨率预处理使准确率达79.37UA/79.68WA,优于以往模型
  • 多窗长STFT数据增强有效缓解数据不足,最高达80.28UA/80.46WA

语音情感识别(SER)通过计算机模型对语音中的情感进行分类。近年来,随着深度学习的发展,SER性能稳步提升,但训练数据仍显不足,导致神经网络过拟合,性能下降。因此,有效的预处理方法与高效利用参数的模型结构至关重要。本文通过八种不同频率-时间分辨率的数据集版本,搜索最优预处理方法;提出一种含高效通道注意力(ECA)的6层卷积神经网络(CNN),ECA模块仅用少量参数即可增强通道特征表示。在IEMOCAP数据集上,提高预处理时的频率分辨率可提升识别性能;在深层卷积后加入ECA能有效增强特征表达。最终取得79.37UA/79.68WA的最佳结果,超过先前模型。此外,为弥补情感语音数据稀缺,实验采用多种短时傅里叶变换(STFT)预处理数据增强策略,从同一样本生成多尺度窗口数据,实现更高性能,达到80.28UA/80.46WA。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) classifies human emotions in speech with a computer model. Recently, performance in SER has steadily increased as deep learning techniques have adapted. However, unlike many domains that use speech data, data for training in the SER model is insufficient. This causes overfitting of training of the neural network, resulting in performance degradation. In fact, successful emotion recognition requires an effective preprocessing method and a model structure that efficiently uses the number of weight parameters. In this study, we propose using eight dataset versions with different frequency-time resolutions to search for an effective emotional speech preprocessing method. We propose a 6-layer convolutional neural network (CNN) model with efficient channel attention (ECA) to pursue an efficient model structure. In particular, the well-positioned ECA blocks can improve channel feature representation with only a few parameters. With the interactive emotional dyadic motion capture (IEMOCAP) dataset, increasing the frequency resolution in preprocessing emotional speech can improve emotion recognition performance. Also, ECA after the deep convolution layer can effectively increase channel feature representation. Consequently, the best result (79.37UA 79.68WA) can be obtained, exceeding the performance of previous SER models. Furthermore, to compensate for the lack of emotional speech data, we experiment with multiple preprocessing data methods that augment trainable data preprocessed with all different settings from one sample. In the experiment, we can achieve the highest result (80.28UA 80.46WA).

语音识别情感分析CNN数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。