分离情绪特征并对齐嵌入,提升嘈杂环境下的语音情感识别能力
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
- 先解耦情绪特征,再跨块对齐嵌入,增强表征鲁棒性
- 在噪声和跨语料数据上准确率显著提升,验证方法有效性
- 适合需要抗噪与跨数据集泛化的语音情感识别场景
真实场景中语音情感识别的性能常受噪声环境和数据集差异影响。本文提出两步方法以提升模型的鲁棒性与泛化能力:首先采用EDRL(情绪解耦表示学习)提取类别特异性判别特征,同时保留情绪类别间的共性;随后通过MEA(多块嵌入对齐)将表示投影到联合判别潜在子空间,最大化其与原始语音输入的协方差。所学得的EDRL-MEA嵌入用于在公开数据集的干净样本上训练情感分类器,并在未见的噪声及跨语料语音样本上评估。该方法在复杂条件下表现优异,证明了其有效性。
原文摘要 · Abstract (English)
Effectiveness of speech emotion recognition in real-world scenarios is often hindered by noisy environments and variability across datasets. This paper introduces a two-step approach to enhance the robustness and generalization of speech emotion recognition models through improved representation learning. First, our model employs EDRL (Emotion-Disentangled Representation Learning) to extract class-specific discriminative features while preserving shared similarities across emotion categories. Next, MEA (Multiblock Embedding Alignment) refines these representations by projecting them into a joint discriminative latent subspace that maximizes covariance with the original speech input. The learned EDRL-MEA embeddings are subsequently used to train an emotion classifier using clean samples from publicly available datasets, and are evaluated on unseen noisy and cross-corpus speech samples. Improved performance under these challenging conditions demonstrates the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。