针对视频表情识别中的标签模糊与类别不平衡问题,提出自适应加权方法提升准确率。
V-NAW: Video-based Noise-aware Adaptive Weighting for Facial Expression Recognition
- 根据帧重要性自适应分配权重,缓解标签模糊影响。
- 在Aff-Wild2数据集上提升视频表情识别准确率,显著优于基线模型。
- 适合关注视频时序建模与弱标签学习的研究者参考。
面部表情识别(FER)在情感分析中至关重要,广泛应用于人机交互与心理评估等任务。本文聚焦于第8届情感行为分析野外挑战赛(ABAW)中的视频表情识别任务,该任务基于Aff-Wild2数据集。现有方法常受标签模糊与类别不平衡影响,导致性能下降。为此,我们提出视频噪声感知自适应加权方法(V-NAW),通过动态调整视频片段中各帧的权重,有效应对标签不一致问题并捕捉表情的时序变化。同时,设计一种简单有效的数据增强策略,降低连续帧间的冗余性,缓解过拟合。大量实验验证了该方法的有效性,在视频表情识别任务中实现显著性能提升。
原文摘要 · Abstract (English)
Facial Expression Recognition (FER) plays a crucial role in human affective analysis and has been widely applied in computer vision tasks such as human-computer interaction and psychological assessment. The 8th Affective Behavior Analysis in-the-Wild (ABAW) Challenge aims to assess human emotions using the video-based Aff-Wild2 dataset. This challenge includes various tasks, including the video-based EXPR recognition track, which is our primary focus. In this paper, we demonstrate that addressing label ambiguity and class imbalance, which are known to cause performance degradation, can lead to meaningful performance improvements. Specifically, we propose Video-based Noise-aware Adaptive Weighting (V-NAW), which adaptively assigns importance to each frame in a clip to address label ambiguity and effectively capture temporal variations in facial expressions. Furthermore, we introduce a simple and effective augmentation strategy to reduce redundancy between consecutive frames, which is a primary cause of overfitting. Through extensive experiments, we validate the effectiveness of our approach, demonstrating significant improvements in video-based FER performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。