arXiv:2503.00202cs.CV2025-03中稿 · WACV2024被引 8

用软标签混合模糊表情数据,提升动态表情识别准确率。

MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression Recognition

  • 通过凸组合视频帧与软标签实现数据增强
  • 在DFEW数据集上超越现有最先进方法
  • 特别适合处理真实场景中的模糊表情

动态面部表情识别(DFER)是计算机视觉的重要任务。为在实际中应用自动DFER,需准确识别常出现在真实数据中的模糊表情。本文提出MIDAS,一种针对DFER的数据增强方法,通过将模糊表情数据与包含多类情绪概率的软标签进行混合。MIDAS通过凸组合视频帧及其对应的情绪类别标签来扩充训练数据,可视为mixup在软标签视频数据上的扩展。该简单扩展在处理模糊表情数据的DFER中表现显著有效。为评估MIDAS,我们在DFEW数据集上进行了实验,结果表明,使用MIDAS增强数据训练的模型性能优于在原始数据上训练的现有最先进方法。

原文摘要 · Abstract (English)

Dynamic facial expression recognition (DFER) is an important task in the field of computer vision. To apply automatic DFER in practice, it is necessary to accurately recognize ambiguous facial expressions, which often appear in data in the wild. In this paper, we propose MIDAS, a data augmentation method for DFER, which augments ambiguous facial expression data with soft labels consisting of probabilities for multiple emotion classes. In MIDAS, the training data are augmented by convexly combining pairs of video frames and their corresponding emotion class labels, which can also be regarded as an extension of mixup to soft-labeled video data. This simple extension is remarkably effective in DFER with ambiguous facial expression data. To evaluate MIDAS, we conducted experiments on the DFEW dataset. The results demonstrate that the model trained on the data augmented by MIDAS outperforms the existing state-of-the-art method trained on the original dataset.

表情识别数据增强软标签视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。