arXiv:2503.06805cs.CVcs.SD2025-03被引 10

融合文本、语音、表情与视频的多模态情感分析方法

Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts

  • 用四个预训练模型融合文本、语音、面部表情和视频特征
  • 情感识别准确率达66.36%,情绪分析达72.15%
  • 适合真实多人对话场景下的情感理解研究

情感识别与情绪分析在语音与语言处理中至关重要,尤其在涉及多方对话的真实场景中。本文提出一种多模态方法,在知名数据集上应对这些挑战。系统整合四种关键模态:使用RoBERTa处理文本,Wav2Vec2处理语音,自研FacialNet分析面部表情,以及从零训练的CNN+Transformer架构处理视频。各模态的特征嵌入被拼接形成多模态向量,用于预测情绪与情感标签。相比单模态方法,该多模态系统表现更优,情感识别准确率为66.36%,情绪分析准确率为72.15%。

原文摘要 · Abstract (English)

Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach to tackle these challenges on a well-known dataset. We propose a system that integrates four key modalities/channels using pre-trained models: RoBERTa for text, Wav2Vec2 for speech, a proposed FacialNet for facial expressions, and a CNN+Transformer architecture trained from scratch for video analysis. Feature embeddings from each modality are concatenated to form a multimodal vector, which is then used to predict emotion and sentiment labels. The multimodal system demonstrates superior performance compared to unimodal approaches, achieving an accuracy of 66.36% for emotion recognition and 72.15% for sentiment analysis.

多模态情感识别对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。