arXiv:2508.02448cs.SDeess.AS2025-08被引 1

15年深度学习语音情感识别进展实证研究,发现性能提升已趋平缓。

Charting 15 years of progress in deep learning for speech emotion recognition: A replication study

  • 对比2009至今的音频与文本模型,系统评估演进效果
  • 发现变压器架构后性能提升出现边际递减并趋于饱和
  • 揭示研究者对进展的认知受模型选择影响,警示未来方向

语音情感识别(SER)长期受益于深度学习方法的应用。更深层、更多参数的模型通常被视为更优,这引出关键问题:现代深度神经网络相比早期版本究竟提升了多少?更重要的是,如何推进未来研究仍至关重要。本文量化分析自2009年首届INTERSPEECH情感挑战赛以来15年的进展,大规模考察了基于语音输入的音频模型与仅依赖转录文本的文本模型。结果表明,在近期引入变压器架构后,性能提升出现边际递减并趋于平台期。此外,我们证明了人们对进展的认知高度依赖于所比较的模型组合。这些发现对当前SER研究的最前沿状态及未来路径具有重要启示。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) has long benefited from the adoption of deep learning methodologies. Deeper models -- with more layers and more trainable parameters -- are generally perceived as being `better' by the SER community. This raises the question -- \emph{how much better} are modern-era deep neural networks compared to their earlier iterations? Beyond that, the more important question of how to move forward remains as poignant as ever. SER is far from a solved problem; therefore, identifying the most prominent avenues of future research is of paramount importance. In the present contribution, we attempt a quantification of progress in the 15 years of research beginning with the introduction of the landmark 2009 INTERSPEECH Emotion Challenge. We conduct a large scale investigation of model architectures, spanning both audio-based models that rely on speech inputs and text-baed models that rely solely on transcriptions. Our results point towards diminishing returns and a plateau after the recent introduction of transformer architectures. Moreover, we demonstrate how perceptions of progress are conditioned on the particular selection of models that are compared. Our findings have important repercussions about the state-of-the-art in SER research and the paths forward

语音识别情感分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。