用统一架构融合脑电与表情视频,提升情绪评估准确率
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment

- 设计共享注意力结构,统一处理脑电与视频信号
- 融合后测试准确率达70.07%,控制时长后仍达62.71%
- 揭示数据时长差异可能影响分类结果,提醒需严格控制
自动情绪评估可受益于神经与行为信号的结合,但多数多模态方法在融合前使用独立的模态特异性特征提取流程。本文提出MUPA²E,一种统一感知框架,通过单一共享的非对称注意力主干网络处理面部视频与脑电(EEG)信号。面部视频以轴折叠帧标记表示,EEG则作为原始多通道波形或投影至空间域用于多模态融合。在DMER数据集上采用分层的受试者无关协议评估,对比单模态视频、单模态EEG及融合视频-EEG配置,含每通道与合并的EEG投影。使用原始记录,较短实验段通过零填充匹配最长时长,步长为30的合并融合在验证集表现最优,测试准确率达70.07%。进一步分析发现,情感类别间记录时长分布不均,导致填充模式可能成为分类线索。通过将所有记录裁剪至20秒共同时长,得到测试准确率62.71%,提供更严格的时长控制评估,消除了时长差异这一潜在分类线索。研究证明了在紧凑统一架构中处理结构差异大的神经与视觉信号的可行性,同时强调了在情感数据集中控制时长相关线索的重要性。
原文摘要 · Abstract (English)
Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion. This paper presents MUPA\textsuperscript{2}E, a unified perception framework that processes facial video and electroencephalography (EEG) through a single shared asymmetric-attention backbone. Facial video is represented through axis-folded frame tokens, while EEG is processed either as a raw multichannel waveform or projected into the spatial domain for multimodal fusion. The framework is evaluated on the DMER dataset under a stratified subject-independent protocol, comparing unimodal video, unimodal EEG, and fused video--EEG configurations with per-channel and merged EEG projections. Using the original recordings, with shorter trials zero-padded to match the longest duration, merged fusion at stride~$30$ achieves the highest validation performance and a test accuracy of $70.07\%$. Further analysis revealed that recording duration is unevenly distributed across the affective classes, making the padding pattern a potential classification cue. Controlling for this factor by cropping all recordings to a common duration of $20$ seconds yielded a test accuracy of $62.71\%$, providing a stricter duration-controlled assessment of the framework in which differences in recording length are removed as a potential classification cue. These findings demonstrate the feasibility of processing structurally different neural and visual signals within a compact unified architecture while highlighting the importance of controlling duration-related cues in affective datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。