现有情绪识别模型难以理解合成语音,因生成过程导致语义与情感特征不匹配。
On the Emotion Understanding of Synthesized Speech

- 通过多数据集、多模型对比,检验合成语音的情绪识别能力
- 发现合成语音与真人语音存在表征差异,导致模型泛化失败
- 指出当前模型依赖文本语义而非语音韵律,影响真实情感理解
情感是语音交互中的核心副语言特征。普遍认为情感识别模型能学习到可迁移的基础表征,因此可用于评估语音合成中的情感表现力。本文系统评估了在不同数据集、判别式与生成式情感识别模型以及多种合成模型上,对合成语音的情感识别表现。结果表明,现有情感识别模型无法有效泛化到合成语音,主要原因是语音生成过程中的话语标记预测导致合成语音与真人语音的表征不一致。此外,生成式语音语言模型倾向于从文本语义推断情感,而忽略副语言线索。总体而言,现有情感识别模型往往依赖非鲁棒性捷径,未能捕捉基础特征,表明语音语言模型中的副语言理解仍具挑战性。
原文摘要 · Abstract (English)
Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible reward or evaluation metric for assessing emotional expressiveness in speech synthesis. In this work, we critically examine this assumption by systematically evaluating Speech Emotion Recognition (SER) on synthesized speech across datasets, discriminative and generative SER models, and diverse synthesis models. We find that current SER models can not generalize to synthesized speech, largely because speech token prediction during synthesis induces a representation mismatch between synthesized and human speech. Moreover, generative Speech Language Models (SLMs) tend to infer emotion from textual semantics while ignoring paralinguistic cues. Overall, our findings suggest that existing SER models often exploit non-robust shortcuts rather than capturing fundamental features, and paralinguistic understanding in SLMs remains challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。