arXiv:2512.16485cs.CVcs.AI2025-12中稿 · TMM被引 3

用眼动行为补足面部表情,提升真实情绪识别准确率

Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors

  • 引入眼动数据构建多模态情绪数据集,捕捉真实情绪
  • 新模型在多任务框架下显著提升情绪识别效果
  • 适合关注真实情绪识别与多模态融合的研究者

情绪识别(ER)旨在从感知数据中分析和识别人类情绪。当前研究主要依赖面部表情识别(FER),但面部表情常被用作社交工具而非真实情绪的体现。为弥合二者差距,本文引入眼动行为作为重要情绪线索,构建了眼动辅助的多模态情绪识别(EMER)数据集。通过自发情绪诱导范式与刺激材料,非侵入式采集眼动序列与注视图,并同步获取面部表情视频。针对多模态与单模态分别标注多视角情绪标签。基于该数据集,设计了简单有效的眼动辅助多模态情绪识别变换器(EMERT),通过模态对抗特征解耦与多任务Transformer,将眼动行为作为面部表情的有力补充。实验采用七种基准评估协议,结果表明EMERT显著优于现有先进方法,验证了眼动建模对鲁棒情绪识别的重要性。本文提供全面分析,推动解决面部表情与真实情绪间的鸿沟。数据集与模型将在https://github.com/kejun1/EMER公开。

原文摘要 · Abstract (English)

Emotion Recognition (ER) is the process of analyzing and identifying human emotions from sensing data. Currently, the field heavily relies on facial expression recognition (FER) because visual channel conveys rich emotional cues. However, facial expressions are often used as social tools rather than manifestations of genuine inner emotions. To understand and bridge this gap between FER and ER, we introduce eye behaviors as an important emotional cue and construct an Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset. To collect data with genuine emotions, spontaneous emotion induction paradigm is exploited with stimulus material, during which non-invasive eye behavior data, like eye movement sequences and eye fixation maps, is captured together with facial expression videos. To better illustrate the gap between ER and FER, multi-view emotion labels for mutimodal ER and FER are separately annotated. Furthermore, based on the new dataset, we design a simple yet effective Eye-behavior-aided MER Transformer (EMERT) that enhances ER by bridging the emotion gap. EMERT leverages modality-adversarial feature decoupling and a multitask Transformer to model eye behaviors as a strong complement to facial expressions. In the experiment, we introduce seven multimodal benchmark protocols for a variety of comprehensive evaluations of the EMER dataset. The results show that the EMERT outperforms other state-of-the-art multimodal methods by a great margin, revealing the importance of modeling eye behaviors for robust ER. To sum up, we provide a comprehensive analysis of the importance of eye behaviors in ER, advancing the study on addressing the gap between FER and ER for more robust ER performance. Our EMER dataset and the trained EMERT models will be publicly available at https://github.com/kejun1/EMER.

情绪识别多模态眼动分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。