测试语音模型在情感矛盾语音中的表现,发现其主要依赖文本而非语音情绪。
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
- 用情感矛盾语音测试四个语音模型的识别能力
- 模型准确率仅比随机猜测略高,依赖文本语义而非语音表达
- 适合关注多模态融合、模型可解释性的研究者
语音语言模型(SLMs)通过联合学习文本与音频表征,旨在实现通用语音理解。尽管取得进展,其泛化能力及音频与文本模态在内部表示中的融合程度仍存争议。本文在情感矛盾语音数据集上评估四个SLMs的情感识别性能,该数据集中语义内容传达一种情绪,而语音表达传达另一种。结果表明,模型主要依赖文本语义而非语音情绪进行判断,文本表征显著主导音频表征。研究发布代码与名为EMIS的合成情感矛盾语音数据集。
原文摘要 · Abstract (English)
Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks. Although promising results have been achieved, there is growing discussion regarding these models' generalization capabilities and the extent to which they truly integrate audio and text modalities in their internal representations. In this work, we evaluate four SLMs on the task of speech emotion recognition using a dataset of emotionally incongruent speech samples, a condition under which the semantic content of the spoken utterance conveys one emotion while speech expressiveness conveys another. Our results indicate that SLMs rely predominantly on textual semantics rather than speech emotion to perform the task, indicating that text-related representations largely dominate over acoustic representations. We release both the code and the Emotionally Incongruent Synthetic Speech dataset (EMIS) to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。