用自监督特征端到端融合语音情绪识别与语音活动检测
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
- 用自监督特征统一建模语音活动检测与情绪识别
- 在IEMOCAP数据集上情绪识别准确率显著提升
- 适合语音识别与情感分析联合优化场景
语音情绪识别(SER)通常依赖语音活动检测(VAD)模型提取语音片段。然而,现有VAD在嘈杂环境下常输出错误语音段,导致后续SER性能下降。为此,本文提出一种基于自监督学习(SSL)特征的端到端(E2E)联合方法,将VAD与SER模块结合:先由VAD模块处理输入的SSL特征,再将分割后的特征送入SER模块,两者联合训练以优化情绪识别效果。在IEMOCAP数据集上的实验表明,该方法显著提升了情绪识别性能。为进一步分析其影响,还对VAD输出及SSL编码器各层权重进行了深入分析。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded performance of subsequent SER models. To address this issue, we propose an end-to-end (E2E) method that integrates VAD and SER using Self-Supervised Learning (SSL) features. The VAD module first receives the SSL features as input, and the segmented SSL features are then fed into the SER module. Both the VAD and SER modules are jointly trained to optimize SER performance. Experimental results on the IEMOCAP dataset demonstrate that our proposed method improves SER performance. Furthermore, to investigate the effect of our proposed method on the VAD and SSL modules, we present an analysis of the VAD outputs and the weights of each layer of the SSL encoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。