arXiv:2503.09929cs.CV2025-03被引 13

用CLIP+TCN提升连续情绪识别准确率,适合做情感计算研究者参考。

Emotion Recognition with CLIP and Sequential Learning

  • 基于CLIP模型在aff-wild2数据集上微调,获得更强视觉特征提取能力。
  • 融合TCN与Transformer结构,在VA、表情和动作单元检测任务中表现更优。
  • 方法简洁高效,适合实际应用中的连续情绪分析场景。

人类情绪识别在人机无缝交互中具有关键作用。本文针对第八届野外情感行为分析研讨会(ABAW)的连续情绪估计、表情识别与动作单元检测挑战,提出一种新方法。通过在aff-wild2数据集上微调CLIP模型,获得具备标注表情标签的视觉特征提取器,显著增强模型鲁棒性。为进一步提升连续情绪识别性能,系统架构中引入时间卷积网络(TCN)与Transformer编码器模块。该集成结构使模型在多项任务上超越基线,展现出更高的识别准确率与效率。

原文摘要 · Abstract (English)

Human emotion recognition plays a crucial role in facilitating seamless interactions between humans and computers. In this paper, we present our innovative methodology for tackling the Valence-Arousal (VA) Estimation Challenge, the Expression Recognition Challenge, and the Action Unit (AU) Detection Challenge, all within the framework of the 8th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW). Our approach introduces a novel framework aimed at enhancing continuous emotion recognition. This is achieved by fine-tuning the CLIP model with the aff-wild2 dataset, which provides annotated expression labels. The result is a fine-tuned model that serves as an efficient visual feature extractor, significantly improving its robustness. To further boost the performance of continuous emotion recognition, we incorporate Temporal Convolutional Network (TCN) modules alongside Transformer Encoder modules into our system architecture. The integration of these advanced components allows our model to outperform baseline performance, demonstrating its ability to recognize human emotions with greater accuracy and efficiency.

情绪识别CLIPTCN多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。