用多语言教师模型蒸馏知识,训练出能识别英法芬语情绪的统一模型。
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
- 通过多教师蒸馏,将英法芬三语教师知识融合到单一学生模型中。
- 在英语数据上加权召回率达72.9,芬兰语未加权召回达63.4,领先基线。
- 特别提升悲伤与中性情绪识别,适合多语言情感交互系统开发。
语音情绪识别(SER)对提升人机交互至关重要。尽管单语言SER已取得进展,但构建多语言系统仍具挑战。本文目标是通过从多个单语言教师模型中蒸馏知识,训练一个能处理多语言的统一学生模型。为此,提出一种新颖的语言感知多教师知识蒸馏方法,基于Wav2Vec2.0构建英、法、芬三语教师模型,并将其知识迁移至单一多语言学生模型。该学生模型表现优异:在英语数据集上加权召回率达72.9,在芬兰语数据集上未加权召回率达63.4,优于微调和传统蒸馏基线。方法在识别悲伤与中性情绪方面尤为突出,但在愤怒和快乐情绪识别上仍有不足。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single model capable of multilingual SER by distilling knowledge from multiple teacher models. To address this, we introduce a novel language-aware multi-teacher knowledge distillation method to advance SER in English, Finnish, and French. It leverages Wav2Vec2.0 as the foundation of monolingual teacher models and then distills their knowledge into a single multilingual student model. The student model demonstrates state-of-the-art performance, with a weighted recall of 72.9 on the English dataset and an unweighted recall of 63.4 on the Finnish dataset, surpassing fine-tuning and knowledge distillation baselines. Our method excels in improving recall for sad and neutral emotions, although it still faces challenges in recognizing anger and happiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。