用肌电与下脸视频融合,提升头显遮挡下的虚拟现实情绪识别
Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

- 融合下脸视频与上脸肌电信号,克服头显遮挡问题
- 在独立受试者测试中达51%宏平均F1,优于纯视觉或纯肌电方法
- 适合做虚拟现实中的情感自适应系统,如心理治疗和沟通训练
头戴式显示器(HMD)会遮挡面部上半部分,导致传统基于图像的面部表情分析不完整,尤其影响需要实时情感评估的应用。本文通过融合下脸视频与上脸肌电图(EMG)信号,实现七类情绪分类(六种基本情绪加中性)。构建了来自20名参与者、配对下脸视频与七通道上脸EMG的同步多模态数据集,使用情绪诱发刺激。在受试者独立测试下,所提晚期融合架构将卷积视觉嵌入与RBF核肌电表示结合,达到51%宏平均F1,优于仅用图像(41%)和仅用肌电(43%)的基线。结果表明,上脸肌电在视觉遮挡下提供鲁棒的补充信息,为自然情境下的虚拟现实多模态情绪识别奠定基础。该方法可用于情感自适应应用,如沟通训练和治疗干预。数据集将在伦理协议下按需共享。
原文摘要 · Abstract (English)
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。