通过三步法提升多人对话情绪识别准确率
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion

- 用音视频同步识别说话人,解决混淆问题
- 文本情感知识迁移到音频视频模态,提升理解力
- 分层融合+复合损失,改善少数情绪类别识别
多人对话中的情绪识别面临说话人混淆和严重类别不平衡的挑战。本文提出一种新框架,包含三项关键创新:(1) 利用音视频同步的说话人识别模块,精准定位当前发言者;(2) 采用知识蒸馏策略,将文本模态中更优的情绪理解能力迁移至音频和视觉模态;(3) 使用分层注意力融合与复合损失函数,缓解类别不平衡问题。在MELD和IEMOCAP数据集上的全面评估显示,该方法分别取得67.75%和72.44%的加权F1分数,尤其在少数情绪类别上表现显著提升。
原文摘要 · Abstract (English)
Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker identification module that leverages audio-visual synchronization to accurately identify the active speaker, (2) a knowledge distillation strategy that transfers superior textual emotion understanding to audio and visual modalities, and (3) hierarchical attention fusion with composite loss functions to handle class imbalance. Comprehensive evaluations on MELD and IEMOCAP datasets demonstrate superior performance, achieving 67.75% and 72.44% weighted F1 scores respectively, with particularly notable improvements on minority emotion classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。