arXiv:2409.06154cs.CV2024-09被引 14

用静态表情数据提升动态表情识别性能,效果显著。

Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data

  • 构建双模态自监督框架,融合静态与动态表情数据
  • 在多个数据集上实现新纪录,最高准确率达76.68%
  • 适合研究情感计算与多任务学习的学者参考

动态面部表情识别(DFER)通过分析表情随时间的变化来推断情绪,相比仅依赖单张图像的静态表情识别(SFER),能提供更丰富的信息。然而当前DFER方法性能不佳,主要因训练样本远少于SFER。鉴于静态与动态表情存在内在关联,我们提出静态转动态(S4D)框架,利用大量SFER数据辅助DFER。S4D采用共享视觉变换器(ViT)编码器-解码器结构,在图像和视频上进行双模态自监督预训练,获得更好的时空表征。随后在静态与动态表情数据集上进行多任务微调,促进情感信息交互。但标准多任务学习出现负迁移,为此我们设计混合适配器专家(MoAE)模块,有效实现任务特异性知识获取与共性知识提取。大量实验表明,S4D在FERV39K、MAFW和DFEW三个基准上分别达到53.65%、58.44%和76.68%的加权平均召回率(WAR),刷新了现有水平。此外,系统性相关性分析进一步揭示了利用SFER数据的潜力。

原文摘要 · Abstract (English)

Dynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on a single snapshot. This temporal analysis provides richer information and promises greater recognition capability. However, current DFER methods often exhibit unsatisfied performance largely due to fewer training samples compared to SFER. Given the inherent correlation between static and dynamic expressions, we hypothesize that leveraging the abundant SFER data can enhance DFER. To this end, we propose Static-for-Dynamic (S4D), a unified dual-modal learning framework that integrates SFER data as a complementary resource for DFER. Specifically, S4D employs dual-modal self-supervised pre-training on facial images and videos using a shared Vision Transformer (ViT) encoder-decoder architecture, yielding improved spatiotemporal representations. The pre-trained encoder is then fine-tuned on static and dynamic expression datasets in a multi-task learning setup to facilitate emotional information interaction. Unfortunately, vanilla multi-task learning in our study results in negative transfer. To address this, we propose an innovative Mixture of Adapter Experts (MoAE) module that facilitates task-specific knowledge acquisition while effectively extracting shared knowledge from both static and dynamic expression data. Extensive experiments demonstrate that S4D achieves a deeper understanding of DFER, setting new state-of-the-art performance on FERV39K, MAFW, and DFEW benchmarks, with weighted average recall (WAR) of 53.65\%, 58.44\%, and 76.68\%, respectively. Additionally, a systematic correlation analysis between SFER and DFER tasks is presented, which further elucidates the potential benefits of leveraging SFER.

表情识别多任务学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。