提出新方法区分难样本与噪声样本,提升视频表情识别鲁棒性。
Robust Dynamic Facial Expression Recognition
- 通过视频片段预测一致性判断样本难易与噪声程度
- 在DFEW和FERV39K上优于当前最先进方法
- 适合研究动态表情识别与抗噪学习的学者
动态面部表情识别(DFER)是基于视频数据自动识别面部表情的新兴研究领域。现有研究多关注噪声与难样本的表示学习,但两者共存问题尚未解决。本文提出一种鲁棒方法,通过评估模型对视频不同采样片段的预测一致性,区分难样本与噪声样本,并据此强化难样本学习、抑制噪声影响。进一步提出关键表达重采样框架与双流分层网络,构建鲁棒动态面部表情识别(RDFER)模型。该框架可识别视频中的主要表情,缓解非目标表情带来的干扰;双序列模型分别捕捉短期面部动作与长期情绪变化。在DFEW和FERV39K等基准数据集上的大量实验表明,RDFER性能超越现有最先进方法。全面分析揭示了预测一致性机制的有效性。本工作对动态表情识别及噪声一致鲁棒学习具有重要意义。代码已公开于[https://github.com/Cross-Innovation-Lab/RDFER]。
原文摘要 · Abstract (English)
The study of Dynamic Facial Expression Recognition (DFER) is a nascent field of research that involves the automated recognition of facial expressions in video data. Although existing research has primarily focused on learning representations under noisy and hard samples, the issue of the coexistence of both types of samples remains unresolved. In order to overcome this challenge, this paper proposes a robust method of distinguishing between hard and noisy samples. This is achieved by evaluating the prediction agreement of the model on different sampled clips of the video. Subsequently, methodologies that reinforce the learning of hard samples and mitigate the impact of noisy samples can be employed. Moreover, to identify the principal expression in a video and enhance the model's capacity for representation learning, comprising a key expression re-sampling framework and a dual-stream hierarchical network is proposed, namely Robust Dynamic Facial Expression Recognition (RDFER). The key expression re-sampling framework is designed to identify the key expression, thereby mitigating the potential confusion caused by non-target expressions. RDFER employs two sequence models with the objective of disentangling short-term facial movements and long-term emotional changes. The proposed method has been shown to outperform current State-Of-The-Art approaches in DFER through extensive experimentation on benchmark datasets such as DFEW and FERV39K. A comprehensive analysis provides valuable insights and observations regarding the proposed agreement. This work has significant implications for the field of dynamic facial expression recognition and promotes the further development of the field of noise-consistent robust learning in dynamic facial expression recognition. The code is available from [https://github.com/Cross-Innovation-Lab/RDFER].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。