arXiv:2409.09511cs.SDcs.AI2024-09被引 4

用可解释声学特征揭示深度语音情感嵌入的内在机制

Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features

  • 通过筛选关键嵌入维度预测声学特征,反推其重要性
  • 能量、频率、频谱、时序特征对情感识别贡献递减
  • 适合关注模型可信度与声学机理的研究者

预训练的深度学习嵌入在语音情感识别(SER)中持续优于人工设计的声学特征,但缺乏明确的可解释性。本文提出一种改进的探针方法,从(i)完整嵌入集和(ii)针对每种情感筛选出的关键嵌入维度中,预测可解释的声学特征(如基频f0、响度)。若关键维度在预测特定情感及对应声学特征上表现更优,则说明这些特征对模型任务至关重要。基于WavLM嵌入与eGeMAPS声学特征,在RAVDESS和SAVEE数据集上验证,结果显示能量、频率、频谱、时序类特征对情感识别的信息贡献呈递减顺序,证明了该探针方法的有效性。

原文摘要 · Abstract (English)

Pre-trained deep learning embeddings have consistently shown superior performance over handcrafted acoustic features in speech emotion recognition (SER). However, unlike acoustic features with clear physical meaning, these embeddings lack clear interpretability. Explaining these embeddings is crucial for building trust in healthcare and security applications and advancing the scientific understanding of the acoustic information that is encoded in them. This paper proposes a modified probing approach to explain deep learning embeddings in the SER space. We predict interpretable acoustic features (e.g., f0, loudness) from (i) the complete set of embeddings and (ii) a subset of the embedding dimensions identified as most important for predicting each emotion. If the subset of the most important dimensions better predicts a given emotion than all dimensions and also predicts specific acoustic features more accurately, we infer those acoustic features are important for the embedding model for the given task. We conducted experiments using the WavLM embeddings and eGeMAPS acoustic features as audio representations, applying our method to the RAVDESS and SAVEE emotional speech datasets. Based on this evaluation, we demonstrate that Energy, Frequency, Spectral, and Temporal categories of acoustic features provide diminishing information to SER in that order, demonstrating the utility of the probing classifier method to relate embeddings to interpretable acoustic features.

语音情感可解释性嵌入分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。