arXiv:2503.17453cs.CV2025-03被引 2

融合ViT与ResNet特征,提升复杂场景下复合情绪识别准确率

Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition

  • 采用ViT与ResNet双分支提取视觉特征并融合
  • 在C-EXPR-DB上达到83.6%准确率,优于单一模型
  • 适合需要高鲁棒性多模态情绪分析的系统开发

本文介绍了在第八届情感行为分析野外(ABAW)竞赛中的成果。多模态情绪识别在情感计算与人机交互中具有重要应用,但在真实场景中,复合情绪识别面临更高的不确定性与模态冲突问题。针对复合表情(CE)识别挑战,本文提出一种融合视觉变压器(ViT)与残差网络(ResNet)特征的多模态情绪识别方法。在C-EXPR-DB和MELD数据集上进行实验,结果表明,在视觉与音频线索复杂的场景(如C-EXPR-DB)中,融合ViT与ResNet特征的模型表现更优,准确率达到83.6%。代码已开源:https://github.com/MyGitHub-ax/8th_ABAW。

原文摘要 · Abstract (English)

This article presents our results for the eighth Affective Behavior Analysis in-the-wild (ABAW) competition.Multimodal emotion recognition (ER) has important applications in affective computing and human-computer interaction. However, in the real world, compound emotion recognition faces greater issues of uncertainty and modal conflicts. For the Compound Expression (CE) Recognition Challenge,this paper proposes a multimodal emotion recognition method that fuses the features of Vision Transformer (ViT) and Residual Network (ResNet). We conducted experiments on the C-EXPR-DB and MELD datasets. The results show that in scenarios with complex visual and audio cues (such as C-EXPR-DB), the model that fuses the features of ViT and ResNet exhibits superior performance.Our code are avalible on https://github.com/MyGitHub-ax/8th_ABAW

情绪识别多模态ViTResNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。