用跨模态提示调优融合面部与生理信号,提升视频情绪识别泛化能力。
Adaptive Physical-Facial Representation Fusion via Subject-Invariant Cross-Modal Prompt Tuning for Video-Based Emotion Recognition

- 通过时频表示生成互补提示,动态调制面部特征而不破坏预训练模型
- 在MAHNOB-HCI和DEAP数据集上准确率优于基线,跨人泛化能力显著提升
- 适合需要高鲁棒性情绪识别的医疗、人机交互等场景
基于面部视频的情绪识别可实现非接触式情感状态推断。尽管面部表情是常用线索,但无法完全反映内在情感状态。远程光电容积脉搏波(rPPG)提供互补生理信息,但易受噪声和个体差异影响,限制了对未见个体的泛化能力。现有多模态方法虽融合面部与rPPG特征,但融合策略常破坏预训练面部表征,且缺乏显式抑制个体特异性变化的机制。为此,本文提出一种主体无关的跨模态提示调优框架。具体地,将rPPG波形转化为抗噪时频表示(TFR),从中生成模态互补提示,调制冻结的视觉变压器(ViT)中的面部令牌。该设计实现有效跨模态交互的同时,保留预训练主干学习到的可泛化面部表征。此外,在每层ViT中引入解耦共享-特定适配器(DSSA),显式分离共享与个体特异性成分,提升跨主体泛化能力。在MAHNOB-HCI和DEAP基准上的实验表明,所提方法在识别准确率与泛化能力上均持续优于强基线,验证了其在视频情绪识别中的有效性。
原文摘要 · Abstract (English)
Emotion recognition from facial videos enables non-contact inference of human emotional states. Although facial expressions are widely used cues, they cannot fully reflect intrinsic affective states. Remote photoplethysmography (rPPG) provides complementary physiological information, but it is highly susceptible to noise and inter-subject variability, limiting generalization to unseen individuals. Existing multimodal methods combine facial and rPPG features, yet their fusion strategies often disrupt pretrained facial representations and lack explicit mechanisms to suppress subject-specific variations. To address these issues, we propose a subject-invariant cross-modal prompt-tuning framework for video-based emotion recognition. Specifically, rPPG waveforms are transformed into noise-robust time-frequency representations (TFRs), from which modality-complementary prompts are generated to modulate facial tokens within a frozen Vision Transformer (ViT). This design enables effective cross-modal interaction while preserving the generalizable facial representations learned by the pretrained backbone. In addition, we introduce a decoupled shared-specific adapter (DSSA) into each ViT layer to explicitly separate subject-shared and subject-specific components, thereby improving cross-subject generalization. Experiments on the MAHNOB-HCI and DEAP benchmarks demonstrate that the proposed method consistently outperforms strong baselines in both recognition accuracy and generalization ability, highlighting its effectiveness for video-based emotion recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。