融合眼动、性格与情境,提升面对面情绪识别准确率。
Modelling the Interplay of Eye-Tracking Temporal Dynamics and Personality for Emotion Detection in Face-to-Face Settings
- 结合眼动时序、性格特质和刺激情境建模情绪
- 情境线索显著提升感知情绪预测(宏平均F1达0.77)
- 性格特质对真实感受情绪识别提升最大(宏平均F1达0.58)
准确识别人类情绪对自适应人机交互至关重要,但在动态对话场景中仍具挑战。本文提出一种人格感知的多模态框架,融合眼动序列、大五人格特质及上下文刺激线索,预测感知与真实情绪。73名参与者观看包含语音的CREMA-D数据集片段,同时提供眼动信号、人格评估与情绪评分。神经模型捕捉眼动时序动态,并融合特质与刺激信息,在感知情绪预测上优于SVM与文献基准。结果表明:(i) 刺激线索显著提升感知情绪预测(宏平均F1最高达0.77);(ii) 人格特质对真实情绪识别贡献最大(宏平均F1最高达0.58)。研究强调融合生理、特质与情境信息可缓解情绪主观性问题。通过区分感知与真实反应,该方法推动多模态情感计算发展,为个性化、生态有效的情绪感知系统提供新路径。
原文摘要 · Abstract (English)
Accurate recognition of human emotions is critical for adaptive human-computer interaction, yet remains challenging in dynamic, conversation-like settings. This work presents a personality-aware multimodal framework that integrates eye-tracking sequences, Big Five personality traits, and contextual stimulus cues to predict both perceived and felt emotions. Seventy-three participants viewed speech-containing clips from the CREMA-D dataset while providing eye-tracking signals, personality assessments, and emotion ratings. Our neural models captured temporal gaze dynamics and fused them with trait and stimulus information, yielding consistent gains over SVM and literature baselines. Results show that (i) stimulus cues strongly enhance perceived-emotion predictions (macro F1 up to 0.77), while (ii) personality traits provide the largest improvements for felt emotion recognition (macro F1 up to 0.58). These findings highlight the benefit of combining physiological, trait-level, and contextual information to address the inherent subjectivity of emotion. By distinguishing between perceived and felt responses, our approach advances multimodal affective computing and points toward more personalized and ecologically valid emotion-aware systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。