arXiv:2606.06200cs.SDeess.AS2026-06中稿 · Interspeech 2026

通过对比学习与反向攻击,提升跨语言语音情感识别的泛化能力。

Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition

论文配图:Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition
图 1 · 摘自论文原文
  • 融合对比学习与说话人对抗学习,增强跨语言情感对齐。
  • 在零样本跨语言场景下,性能显著优于传统训练策略。
  • 适合无目标语言标注数据的跨语言情感分析任务。

零样本跨语言语音情感识别(SER)因语言间分布差异及目标语言缺乏情感标注而面临挑战。仅在源语言数据上训练的模型在未见目标语言上常出现泛化性能下降。为此,我们提出一种情感判别性表征学习方法,结合监督对比学习与说话人对抗学习。对比学习促进跨语言情感对齐,说话人对抗学习抑制说话人相关特征,鼓励生成说话人无关表征。在零样本跨语言SER设置下的实验表明,所提方法显著优于传统训练策略。

原文摘要 · Abstract (English)

Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language. Under such conditions, models trained solely on source-language data frequently suffer from degraded generalization when evaluated on unseen target languages. To address this limitation, we propose an emotion-discriminative representation learning method that integrates supervised contrastive learning and speaker adversarial learning. The contrastive learning promotes cross-lingual emotion alignment, while speaker adversarial learning suppresses speaker-related cues to encourage speaker-invariant representations. Experimental results under a zero-shot cross-lingual SER setting demonstrate that the proposed method significantly improves SER performance over conventional training strategies.

语音情感识别跨语言零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。