用软标签混合视频帧,提升模糊表情识别准确率
Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation
- 通过软标签线性组合视频帧和标签,扩展mixup到动态表情数据
- 在DFEW和FERV39k-Plus上,模型性能超越现有最先进方法
- 适合需要处理真实场景中模糊情绪识别的研究者
动态面部表情识别(DFER)旨在从面部表情视频序列中估计情绪。在实际应用中,准确识别模糊表情——在自然环境下频繁出现——至关重要。本文提出MIDAS,一种基于软标签的数据增强方法,用于提升模糊表情数据的DFER性能。MIDAS通过凸组合视频帧及其对应的情绪类别标签来扩充训练数据,将mixup推广至软标签视频数据,提供了一种简单而高效处理DFER中模糊性的方法。为评估MIDAS,我们在DFEW数据集和新构建的FERV39k-Plus数据集上进行了实验,该数据集为现有DFER数据集赋予了软标签。结果表明,使用MIDAS增强数据训练的模型,在性能上优于在原始数据集上训练的最先进方法。
原文摘要 · Abstract (English)
Dynamic facial expression recognition (DFER) is a task that estimates emotions from facial expression video sequences. For practical applications, accurately recognizing ambiguous facial expressions -- frequently encountered in in-the-wild data -- is essential. In this study, we propose MIDAS, a data augmentation method designed to enhance DFER performance for ambiguous facial expression data using soft labels representing probabilities of multiple emotion classes. MIDAS augments training data by convexly combining pairs of video frames and their corresponding emotion class labels. This approach extends mixup to soft-labeled video data, offering a simple yet highly effective method for handling ambiguity in DFER. To evaluate MIDAS, we conducted experiments on both the DFEW dataset and FERV39k-Plus, a newly constructed dataset that assigns soft labels to an existing DFER dataset. The results demonstrate that models trained with MIDAS-augmented data achieve superior performance compared to the state-of-the-art method trained on the original dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。