通过眼动行为提升情绪识别准确率,弥补面部表情的局限。
Smile upon the Face but Sadness in the Eyes: Emotion Recognition based on Facial Expressions and Eye Behaviors
- 引入眼动数据构建新型多模态情绪数据集EMER。
- 新模型在情绪识别上显著优于现有方法。
- 适合关注真实情绪建模与人机交互的研究者。
情绪识别(ER)是从给定数据中识别人类情绪的过程。当前领域主要依赖面部表情识别(FER),因面部表情包含丰富情绪线索。然而,面部表情未必准确反映真实情绪,基于FER的结果可能误导ER。为理解并弥合这一差距,我们引入眼动行为作为重要情绪线索,创建了眼动辅助多模态情绪识别(EMER)数据集。不同于现有数据集,EMER采用刺激物诱导的自发情绪生成方法,融合非侵入性眼动数据(如眼动轨迹和注视图)与面部视频,以获取自然且准确的人类情绪。首次在数据集中同时提供ER和FER标注,支持对两者差异的全面分析。此外,我们设计了新架构EMERT,通过模态对抗特征解耦与多任务Transformer,高效建模眼动行为,有效补充面部表情。实验中,我们设置了七种多模态基准评估协议,结果表明EMERT显著优于其他先进方法,凸显眼动建模对鲁棒情绪识别的重要性。
原文摘要 · Abstract (English)
Emotion Recognition (ER) is the process of identifying human emotions from given data. Currently, the field heavily relies on facial expression recognition (FER) because facial expressions contain rich emotional cues. However, it is important to note that facial expressions may not always precisely reflect genuine emotions and FER-based results may yield misleading ER. To understand and bridge this gap between FER and ER, we introduce eye behaviors as an important emotional cues for the creation of a new Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset. Different from existing multimodal ER datasets, the EMER dataset employs a stimulus material-induced spontaneous emotion generation method to integrate non-invasive eye behavior data, like eye movements and eye fixation maps, with facial videos, aiming to obtain natural and accurate human emotions. Notably, for the first time, we provide annotations for both ER and FER in the EMER, enabling a comprehensive analysis to better illustrate the gap between both tasks. Furthermore, we specifically design a new EMERT architecture to concurrently enhance performance in both ER and FER by efficiently identifying and bridging the emotion gap between the two.Specifically, our EMERT employs modality-adversarial feature decoupling and multi-task Transformer to augment the modeling of eye behaviors, thus providing an effective complement to facial expressions. In the experiment, we introduce seven multimodal benchmark protocols for a variety of comprehensive evaluations of the EMER dataset. The results show that the EMERT outperforms other state-of-the-art multimodal methods by a great margin, revealing the importance of modeling eye behaviors for robust ER. To sum up, we provide a comprehensive analysis of the importance of eye behaviors in ER, advancing the study on addressing the gap between FER and ER for more robust ER performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。