用无标签语音数据提升情感与意图识别准确率
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
- 结合少量标注数据与大量无标签数据,采用半监督学习训练端到端模型
- 融合最优模型后,在情感与意图识别上分别提升12.3%和10.4%的平衡指标
- 适合语音交互、智能客服等需低成本标注的场景
从语音中进行情感与意图识别在人机交互中至关重要,且已被广泛研究。社交媒体平台、聊天机器人等技术的发展带来了海量语音数据,但人工标注成本高昂,制约了机器学习模型的训练。为此,本文提出将半监督学习应用于大规模未标注数据与小规模标注数据的联合训练。我们构建端到端的声学与语言模型,均采用多任务学习实现情感与意图识别。比较了FixMatch与FullMatch两种半监督学习方法。实验表明,半监督方法显著提升了语音情感与意图识别性能。最优模型的后期融合结果在联合识别平衡指标上分别优于声学与文本基线12.3%和10.4%。
原文摘要 · Abstract (English)
Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech data streaming from users. Nevertheless, annotating such data manually is expensive, making it challenging to train machine learning models for recognition purposes. To this end, we propose applying semi-supervised learning to incorporate a large scale of unlabelled data alongside a relatively smaller set of labelled data. We train end-to-end acoustic and linguistic models, each employing multi-task learning for emotion and intent recognition. Two semi-supervised learning approaches, including fix-match learning and full-match learning, are compared. The experimental results demonstrate that the semi-supervised learning approaches improve model performance in speech emotion and intent recognition from both acoustic and text data. The late fusion of the best models outperforms the acoustic and text baselines by joint recognition balance metrics of 12.3% and 10.4%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。