用无标签音视频数据提升语音情感识别,减少对标注数据的依赖。
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
- 通过知识蒸馏,让轻量学生模型学习大模型的音频视觉情感特征。
- 在RAVDESS和CREMA-D数据集上,显著降低对大量标注数据的需求。
- 适合资源受限场景下部署高精度语音情感识别系统。
语音接口在人机交互中日益重要,语音情感识别(SER)可依据用户情绪定制响应。人类通过多模态的音频与视觉线索传递情绪,结合两者开发SER系统更有效。然而,大规模标注数据的收集成本高昂。本文提出一种名为LightweightSER(LiSER)的知识蒸馏框架,利用无标签音视频数据,借助基于先进语音与人脸表示模型的大教师模型,将语音情感与面部表情的知识迁移到轻量级学生模型中。在RAVDESS和CREMA-D两个基准数据集上的实验表明,该方法能有效减少对大规模标注数据的依赖。
原文摘要 · Abstract (English)
Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues, developing SER systems using both the modalities is beneficial. However, collecting a vast amount of labeled data for their development is expensive. This paper proposes a knowledge distillation framework called LightweightSER (LiSER) that leverages unlabeled audio-visual data for SER, using large teacher models built on advanced speech and face representation models. LiSER transfers knowledge regarding speech emotions and facial expressions from the teacher models to lightweight student models. Experiments conducted on two benchmark datasets, RAVDESS and CREMA-D, demonstrate that LiSER can reduce the dependence on extensive labeled datasets for SER tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。