arXiv:2602.18802cs.SD2026-02中稿 · ICASSP2026

多通道语音增强显著提升嘈杂环境下的情绪识别准确率

Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition

  • 融合DNN-WPE与掩码MVDR的多通道前端分离目标语音
  • 情绪识别准确率提升最高达9.5%,相对增益17.1%
  • 适用于嘈杂场景下情绪识别,尤其适合跨数据集泛化

本文强调在鸡尾酒会场景中,多通道语音增强(MCSE)对语音情绪识别(ER)的关键作用。采用集成DNN-WPE与基于掩码的MVDR的多通道语音去混响与分离前端,从混合语音中提取目标说话人语音,再输入基于HuBERT和ViT的语音与视觉特征的下游情绪识别模型。在使用IEMOCAP和MSP-FACE数据集构建的混合语音上实验表明,该方法输出的性能持续优于单通道基线:a) 基于Conformer的度量GAN;b) WavLM自监督特征结合可选的语音增强-情绪识别联合微调。在加权准确率、未加权准确率和F1值上分别获得最高9.5%、8.5%和9.1%的绝对提升(相对提升分别为17.1%、14.7%和16.0%)。此外,基于IEMOCAP训练的MCSE前端在零样本迁移至MSP-FACE数据时也表现出良好泛化能力。

原文摘要 · Abstract (English)

This paper highlights the critical importance of multi-channel speech enhancement (MCSE) for speech emotion recognition (ER) in cocktail party scenarios. A multi-channel speech dereverberation and separation front-end integrating DNN-WPE and mask-based MVDR is used to extract the target speaker's speech from the mixture speech, before being fed into the downstream ER back-end using HuBERT- and ViT-based speech and visual features. Experiments on mixture speech constructed using the IEMOCAP and MSP-FACE datasets suggest the MCSE output consistently outperforms domain fine-tuned single-channel speech representations produced by: a) Conformer-based metric GANs; and b) WavLM SSL features with optional SE-ER dual task fine-tuning. Statistically significant increases in weighted, unweighted accuracy and F1 measures by up to 9.5%, 8.5% and 9.1% absolute (17.1%, 14.7% and 16.0% relative) are obtained over the above single-channel baselines. The generalization of IEMOCAP trained MCSE front-ends are also shown when being zero-shot applied to out-of-domain MSP-FACE data.

语音增强情绪识别多通道处理鸡尾酒会

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。