arXiv:2512.18298cs.SDcs.CL2025-12被引 1

融合变压器与卷积网络,提升嘈杂环境下的语音情感识别准确率与可解释性。

Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition

  • 双流架构结合上下文建模与频谱稳定性,增强抗噪能力。
  • 在四个数据集上均达到顶尖准确率,噪声环境下性能显著优于单分支模型。
  • 通过SHAP和Score-CAM提供细粒度可视化解释,揭示模型决策逻辑。

语音情感识别系统在真实环境中常因不可预测的声学干扰而性能下降。同时,深度学习模型的黑箱特性也限制其在高可信度场景的应用。为此,我们提出一种混合变压器-卷积神经网络框架,融合Wav2Vec 2.0的上下文建模能力与一维卷积网络的频谱稳定性。该双流架构处理原始波形以捕捉长时程依赖,同时通过自定义的注意力时间池化机制提取噪声鲁棒的谱特征(MFCC、ZCR、RMSE)。我们在RAVDESS、TESS、SAVEE和CREMA-D四个基准数据集上进行了全面验证,并使用SAS-KIIT数据集中的真实噪声样本施加非平稳声学干扰以严格测试鲁棒性。所提框架在所有数据集上均展现出优越的泛化能力与最先进的准确率,显著优于单分支基线模型。此外,通过引入SHAP和Score-CAM,我们实现了细粒度可视化解释,揭示模型如何在复杂环境噪声中动态调整对时序与谱线索的关注以保持可靠性。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in trust-sensitive applications. To bridge this gap, we propose a Hybrid Transformer-CNN framework that unifies the contextual modeling of Wav2Vec 2.0 with the spectral stability of 1D-Convolutional Neural Networks. Our dual-stream architecture processes raw waveforms to capture long-range temporal dependencies while simultaneously extracting noise-resistant spectral features (MFCC, ZCR, RMSE) via a custom Attentive Temporal Pooling mechanism. We conducted extensive validation across four diverse benchmark datasets: RAVDESS, TESS, SAVEE, and CREMA-D. To rigorously test robustness, we subjected the model to non-stationary acoustic interference using real-world noise profiles from the SAS-KIIT dataset. The proposed framework demonstrates superior generalization and state-of-the-art accuracy across all datasets, significantly outperforming single-branch baselines under realistic environmental interference. Furthermore, we address the ``black-box" problem by integrating SHAP and Score-CAM into the evaluation pipeline. These tools provide granular visual explanations, revealing how the model strategically shifts attention between temporal and spectral cues to maintain reliability in the presence of complex environmental noise.

语音情感识别可解释性抗噪融合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。