融合音视频特征,提升野生环境下情绪识别与犹豫度判断性能
HSEmotion Team at ABAW-8 Competition: Audiovisual Ambivalence/Hesitancy, Emotional Mimicry Intensity and Facial Expression Recognition
- 用预训练模型提取面部表情与语音文本特征,融合多模态信息
- 在三个任务上显著超越基线,验证了方法有效性
- 适合关注跨模态情感分析、真实场景行为识别的研究者
本文汇报了我们在第八届野外情感行为分析(ABAW)竞赛中的成果。我们结合由预训练模型提取的面部情绪描述符(基于自研EmotiEffLib库)、声学特征以及语音识别所得文本嵌入。将帧级特征聚合后输入简单分类器(如单隐藏层前馈神经网络),以预测情绪矛盾/犹豫性及面部表情。针对表情识别任务,我们还利用预训练模型筛选高置信度视频帧,避免其被领域特定视频分类器处理。情感模仿强度的视频级预测通过聚合帧级特征并训练多层感知机实现。三个任务的实验结果表明,本方法在验证集指标上显著优于现有基线。
原文摘要 · Abstract (English)
This article presents our results for the eighth Affective Behavior Analysis in-the-Wild (ABAW) competition. We combine facial emotional descriptors extracted by pre-trained models, namely, our EmotiEffLib library, with acoustic features and embeddings of texts recognized from speech. The frame-level features are aggregated and fed into simple classifiers, e.g., multi-layered perceptron (feed-forward neural network with one hidden layer), to predict ambivalence/hesitancy and facial expressions. In the latter case, we also use the pre-trained facial expression recognition model to select high-score video frames and prevent their processing with a domain-specific video classifier. The video-level prediction of emotional mimicry intensity is implemented by simply aggregating frame-level features and training a multi-layered perceptron. Experimental results for three tasks from the ABAW challenge demonstrate that our approach significantly increases validation metrics compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。