arXiv:2509.05634eess.AScs.CL2025-09中稿 · 13th Conference on…

词汇特征在语音情感识别中表现不输声学特征,甚至更优。

On the Contribution of Lexical Features to Speech Emotion Recognition

  • 用提取的词汇内容做情感识别,不依赖声学信号。
  • 在MELD数据集上词汇模型F1达51.5%,优于更大参数的声学模型。
  • 分析了自监督语音/文本表示与降噪对性能的影响。

尽管副语言线索常被视为语音情感识别(SER)的主要驱动因素,本文研究了从语音中提取的词汇内容的作用,发现其性能可与声学模型媲美,甚至在某些情况下更优。在MELD数据集上,基于词汇的方法取得了51.5%的加权F1分数(WF1),而参数更大的纯声学管道仅达到49.3%。此外,我们分析了不同的自监督(SSL)语音与文本表示,对基于Transformer的编码器进行了逐层研究,并评估了音频去噪的影响。

原文摘要 · Abstract (English)

Although paralinguistic cues are often considered the primary drivers of speech emotion recognition (SER), we investigate the role of lexical content extracted from speech and show that it can achieve competitive and in some cases higher performance compared to acoustic models. On the MELD dataset, our lexical-based approach obtains a weighted F1-score (WF1) of 51.5%, compared to 49.3% for an acoustic-only pipeline with a larger parameter count. Furthermore, we analyze different self-supervised (SSL) speech and text representations, conduct a layer-wise study of transformer-based encoders, and evaluate the effect of audio denoising.

语音情感识别词汇特征自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。