arXiv:2409.16920eess.AScs.AI2024-09中稿 · ICASSP 2025被引 13

对比人类与自监督模型在跨语言语音情感识别中的表现

Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models

  • 用分层分析和高效微调策略比较模型与人类表现
  • 模型经知识迁移后可达到母语者水平,但方言影响显著
  • 揭示了人与模型在不同情绪下的识别行为差异

利用自监督学习(SSL)模型进行语音情感识别(SER)已证明有效,但跨语言场景研究仍有限。本研究对比人类与SSL模型在单语言、跨语言及迁移学习场景下的表现,开展分层分析并探索参数高效微调策略。进一步在话语和片段层面比较模型与人类的SER能力,并通过人类评估研究方言对跨语言SER的影响。结果表明,经过适当知识迁移的模型可适应目标语言,性能接近母语者;缺乏语言和副语言背景的人类对方言影响极为敏感。此外,人类与模型在不同情绪下表现出不同的识别模式。这些发现为SSL模型在跨语言SER中的能力提供了新见解,凸显其与人类情感感知的相似性与差异性。

原文摘要 · Abstract (English)

Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an exploration of parameter-efficient fine-tuning strategies in monolingual, cross-lingual, and transfer learning contexts. We further compare the SER ability of models and humans at both utterance- and segment-levels. Additionally, we investigate the impact of dialect on cross-lingual SER through human evaluation. Our findings reveal that models, with appropriate knowledge transfer, can adapt to the target language and achieve performance comparable to native speakers. We also demonstrate the significant effect of dialect on SER for individuals without prior linguistic and paralinguistic background. Moreover, both humans and models exhibit distinct behaviors across different emotions. These results offer new insights into the cross-lingual SER capabilities of SSL models, underscoring both their similarities to and differences from human emotion perception.

语音情感识别自监督学习跨语言人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。