arXiv:2503.21806cs.CLcs.AI2025-03中稿 · ICME 2025被引 5

用对比学习提升跨语言语音情绪识别,零样本效果显著。

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

  • 两阶段训练对齐语音与语言特征,生成情感感知且跨语言通用的表示。
  • 在多语言数据集上实现零样本情绪识别,新语言和新数据集均有效。
  • 构建大规模合成数据集M5SER,推动跨语言情绪识别研究。

多语言语音情绪识别旨在通过无接触方式跨语言判断说话者情绪状态。然而,语音特征差异与语言多样性给零样本语音情绪识别带来挑战,尤其在多语言数据集上。本文提出利用对比学习优化多语言语音特征,并扩展大语言模型以实现零样本多语言语音情绪估计。具体地,采用新颖的两阶段训练框架,将语音信号与语言特征在情感空间中对齐,捕捉兼具情感敏感性与语言无关性的语音表征。为推进该领域研究,我们构建了一个大规模合成多语言语音情绪数据集M5SER。实验表明,所提方法在语音情绪识别及零样本跨语言情绪识别上均有效,涵盖此前未见的数据集和语言。

原文摘要 · Abstract (English)

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses significant challenges for zero-shot speech emotion recognition, especially with multilingual datasets. In this paper, we propose leveraging contrastive learning to refine multilingual speech features and extend large language models for zero-shot multilingual speech emotion estimation. Specifically, we employ a novel two-stage training framework to align speech signals with linguistic features in the emotional space, capturing both emotion-aware and language-agnostic speech representations. To advance research in this field, we introduce a large-scale synthetic multilingual speech emotion dataset, M5SER. Our experiments demonstrate the effectiveness of the proposed method in both speech emotion recognition and zero-shot multilingual speech emotion recognition, including previously unseen datasets and languages.

语音识别跨语言对比学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。