arXiv:2409.19585cs.SDcs.CL2024-09中稿 · APSIPA ASC 2024被引 7

通过两阶段框架提升噪声中语音情感识别准确率,尤其针对人声干扰场景。

Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions

  • 先提取目标说话人语音,再进行情感识别,分步处理噪声干扰。
  • 相比基线模型,未加权准确率提升14.33%,有效降低人声噪声影响。
  • 在异性混合语音场景下表现更优,适合实际复杂对话环境应用。

在嘈杂环境下构建鲁棒的语音情感识别(SER)系统面临噪声特性多样的挑战。以往研究大多未考虑人声噪声的影响,限制了SER的实际应用。本文提出一种新型两阶段框架,将目标说话人提取(TSE)与SER相结合:第一阶段训练TSE模型,从语音混合中提取目标说话人语音;第二阶段利用提取后的语音进行SER训练。此外,还探索了在第二阶段联合训练TSE与SER模型的方案。实验结果表明,本系统相比不使用TSE的基线模型,未加权准确率(UA)提升14.33%,验证了该框架在缓解人声噪声影响方面的有效性。进一步实验还考虑了说话人性别因素,发现该框架在异性混合语音场景中表现尤为出色。

原文摘要 · Abstract (English)

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application scope of SER. In this paper, we propose a novel two-stage framework for the problem by cascading target speaker extraction (TSE) method and SER. We first train a TSE model to extract the speech of target speaker from a mixture. Then, in the second stage, we utilize the extracted speech for SER training. Additionally, we explore a joint training of TSE and SER models in the second stage. Our developed system achieves a 14.33% improvement in unweighted accuracy (UA) compared to a baseline without using TSE method, demonstrating the effectiveness of our framework in mitigating the impact of human speech noise. Moreover, we conduct experiments considering speaker gender, showing that our framework performs particularly well in different-gender mixture.

语音情感识别说话人分离噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。