arXiv:2506.02059cs.SDcs.CL2025-06中稿 · Interspeech 2025被引 7

用自监督学习提升低资源语言情感识别效果

Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition

  • 采用对比学习和BYOL方法,无需标注数据训练
  • 在乌尔都语、德语、孟加拉语上分别提升10.6%~15.2%准确率
  • 为少数语言情感识别提供可解释的鲁棒解决方案

语音情感识别(SER)在深度学习推动下取得显著进展,但低资源语言(LRLs)因标注数据稀缺仍面临挑战。本文探索无监督学习以改善低资源环境下的SER。具体而言,研究了对比学习(CL)和自举你自己潜在表示(BYOL)两种自监督方法,以增强跨语言泛化能力。实验显示,该方法在乌尔都语、德语和孟加拉语上的F1分数分别提升10.6%、15.2%和13.9%,证明其在低资源语言中的有效性。此外,通过分析模型行为,揭示影响跨语言性能的关键因素,并指出低资源SER面临的挑战。本工作为构建更具包容性、可解释性和鲁棒性的少数语言情感识别系统奠定基础。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) has seen significant progress with deep learning, yet remains challenging for Low-Resource Languages (LRLs) due to the scarcity of annotated data. In this work, we explore unsupervised learning to improve SER in low-resource settings. Specifically, we investigate contrastive learning (CL) and Bootstrap Your Own Latent (BYOL) as self-supervised approaches to enhance cross-lingual generalization. Our methods achieve notable F1 score improvements of 10.6% in Urdu, 15.2% in German, and 13.9% in Bangla, demonstrating their effectiveness in LRLs. Additionally, we analyze model behavior to provide insights on key factors influencing performance across languages, and also highlighting challenges in low-resource SER. This work provides a foundation for developing more inclusive, explainable, and robust emotion recognition systems for underrepresented languages.

情感识别自监督低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。