arXiv:2511.14969eess.AScs.AI2025-11被引 1

通过身份感知迁移学习提升对话情绪识别质量

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

  • 用身份一致性验证数据,确保音频、文本、人脸对齐
  • 融合语音、人脸与情感调优文本,MELD准确率达64.8%
  • 适合关注多模态情绪识别质量与小众情绪类别的研究者

本文针对对话中多模态情绪识别(MERC)的数据质量问题,提出系统性质量控制与多阶段迁移学习方法。构建了针对MELD与IEMOCAP数据集的质量控制流程,验证说话人身份、音文对齐及人脸检测准确性。基于说话人与人脸识别的迁移学习,假设身份判别嵌入不仅捕捉稳定的声学与面部特征,还包含个体特有的情绪表达模式。采用RecoMadeEasy(R)提取512维说话人与人脸嵌入,微调MPNet-v2获取情感感知文本表示,并通过在单模态数据集上训练的情绪专用MLP适配这些特征。基于MAMBA的三模态融合在MELD上达到64.8%准确率,在IEMOCAP上达74.3%。结果表明,结合身份感知的音视频嵌入与情感调优文本表示,在质量可控数据子集上可实现稳定且具有竞争力的多模态情绪识别性能,为挑战性低频情绪类别提供了改进基础。

原文摘要 · Abstract (English)

This paper addresses data quality issues in multimodal emotion recognition in conversation (MERC) through systematic quality control and multi-stage transfer learning. We implement a quality control pipeline for MELD and IEMOCAP datasets that validates speaker identity, audio-text alignment, and face detection. We leverage transfer learning from speaker and face recognition, assuming that identity-discriminative embeddings capture not only stable acoustic and Facial traits but also person-specific patterns of emotional expression. We employ RecoMadeEasy(R) engines for extracting 512-dimensional speaker and face embeddings, fine-tune MPNet-v2 for emotion-aware text representations, and adapt these features through emotion-specific MLPs trained on unimodal datasets. MAMBA-based trimodal fusion achieves 64.8% accuracy on MELD and 74.3% on IEMOCAP. These results show that combining identity-based audio and visual embeddings with emotion-tuned text representations on a quality-controlled subset of data yields consistent competitive performance for multimodal emotion recognition in conversation and provides a basis for further improvement on challenging, low-frequency emotion classes.

情绪识别多模态迁移学习数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。