arXiv:2503.11935cs.CV2025-03被引 1

用注意力机制和比例融合提升表情识别准确率

Design of an Expression Recognition Solution Based on the Global Channel-Spatial Attention Mechanism and Proportional Criterion Fusion

  • 设计全局通道-空间注意力增强音视频特征
  • 在ABAW竞赛中验证,验证集排名第三
  • 适合多模态情感分析与人机交互研究者

面部表情识别是人机交互领域具有广阔应用前景的挑战性分类任务。本文介绍了将在2025年CVPR会议期间举行的第八届情感与行为分析野外竞赛(8th ABAW Competition)中采用的方法。首先,对原始视频进行频率掩码处理及等时间间隔采样;其次,基于残差混合卷积神经网络和多分支卷积神经网络,分别设计图像与音频序列的特征提取模型,并提出全局通道-空间注意力机制,以增强音视频模态的初始特征表示。最后,采用基于比例准则的决策融合策略,融合双模态分类结果,生成情感概率向量并输出最终情绪分类。此外,设计了粗-细粒度损失函数以优化网络整体性能,显著提升表情识别准确率。在8th ABAW竞赛的面部表情识别任务中,本方法在官方验证集上取得第三名的成绩,充分验证了所提方法的有效性与竞争力。

原文摘要 · Abstract (English)

Facial expression recognition is a challenging classification task that holds broad application prospects in the field of human-computer interaction. This paper aims to introduce the method we will adopt in the 8th Affective and Behavioral Analysis in the Wild (ABAW) Competition, which will be held during the Conference on Computer Vision and Pattern Recognition (CVPR) in 2025.First of all, we apply the frequency masking technique and the method of extracting data at equal time intervals to conduct targeted processing on the original videos. Then, based on the residual hybrid convolutional neural network and the multi-branch convolutional neural network respectively, we design feature extraction models for image and audio sequences. In particular, we propose a global channel-spatial attention mechanism to enhance the features initially extracted from both the audio and image modalities respectively.Finally, we adopt a decision fusion strategy based on the proportional criterion to fuse the classification results of the two single modalities, obtain an emotion probability vector, and output the final emotional classification. We also design a coarse - fine granularity loss function to optimize the performance of the entire network, which effectively improves the accuracy of facial expression recognition.In the facial expression recognition task of the 8th ABAW Competition, our method ranked third on the official validation set. This result fully confirms the effectiveness and competitiveness of the method we have proposed.

表情识别多模态融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。