arXiv:2501.03103cs.CV2025-01ECCV被引 7

融合视频与生理信号,用注意力机制提升长序列情感识别准确率

MVP: Multimodal Emotion Recognition based on Video and Physiological Signals

  • 采用注意力机制处理1-2分钟的长时序视频与生理信号
  • 在面部表情、皮肤电导、心电/脉搏信号上均超越现有方法
  • 适合做多模态情感计算或人机交互研究的开发者

人类情感涉及行为、生理和认知的复杂变化。当前主流模型仍依赖传统机器学习融合行为与生理信号,未充分使用深度学习技术。本文提出MVP架构,专为视频与生理信号的多模态融合设计,通过注意力机制支持长达1-2分钟的输入序列。我们对比了多种视频与生理信号骨干网络,评估了MVP在面部视频、皮肤电导(EDA)及心电(ECG)/脉搏波(PPG)数据上的表现。结果表明,MVP在各类数据上均优于现有情感识别方法。

原文摘要 · Abstract (English)

Human emotions entail a complex set of behavioral, physiological and cognitive changes. Current state-of-the-art models fuse the behavioral and physiological components using classic machine learning, rather than recent deep learning techniques. We propose to fill this gap, designing the Multimodal for Video and Physio (MVP) architecture, streamlined to fuse video and physiological signals. Differently then others approaches, MVP exploits the benefits of attention to enable the use of long input sequences (1-2 minutes). We have studied video and physiological backbones for inputting long sequences and evaluated our method with respect to the state-of-the-art. Our results show that MVP outperforms former methods for emotion recognition based on facial videos, EDA, and ECG/PPG.

情感识别多模态注意力机制生理信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。