arXiv:2512.23076cs.LGcs.AI2025-12中稿 · IEEE Transactions …被引 2

通过高阶相关性建模,提升多模态情感识别的跨被试鲁棒性。

Multimodal Functional Maximum Correlation for Emotion Recognition

  • 用双总相关性目标直接捕捉多模态协同动态,避免依赖成对对比损失。
  • 在三个基准上均达最优或领先性能,如CEAP-360VR中准确率提升至86.8%。
  • 适合关注生理信号融合与自监督学习的情感计算研究者使用。

情绪状态表现为中枢与自主神经系统间协调但异质的生理反应,给情感计算中的多模态表示学习带来根本挑战。由于情感标注稀缺且主观性强,自监督学习(SSL)成为有效解决方案。然而,现有多数方法依赖成对对齐目标,难以刻画超过两模态间的依赖关系,也未能捕捉脑电与自主神经响应的高阶交互。为此,我们提出多模态功能最大相关性(MFMC),一种基于双总相关性(DTC)目标的原理性自监督框架,通过紧致夹心界并利用基于函数最大相关分析(FMCA)的迹代理优化,直接建模联合多模态交互。在三个公开情感计算基准上的实验表明,MFMC在被试内与被试间评估协议下均表现优异:仅使用皮电反应(EDA)信号时,CEAP-360VR的被试内准确率从78.9%提升至86.8%,被试间准确率从27.5%提升至33.1%;在最具有挑战性的MAHNOB-HCI EEG被试间划分中,性能仅比最优方法低0.8个百分点。代码已开源。

原文摘要 · Abstract (English)

Emotional states manifest as coordinated yet heterogeneous physiological responses across central and autonomic systems, posing a fundamental challenge for multimodal representation learning in affective computing. Learning such joint dynamics is further complicated by the scarcity and subjectivity of affective annotations, which motivates the use of self-supervised learning (SSL). However, most existing SSL approaches rely on pairwise alignment objectives, which are insufficient to characterize dependencies among more than two modalities and fail to capture higher-order interactions arising from coordinated brain and autonomic responses. To address this limitation, we propose Multimodal Functional Maximum Correlation (MFMC), a principled SSL framework that maximizes higher-order multimodal dependence through a Dual Total Correlation (DTC) objective. By deriving a tight sandwich bound and optimizing it using a functional maximum correlation analysis (FMCA) based trace surrogate, MFMC captures joint multimodal interactions directly, without relying on pairwise contrastive losses. Experiments on three public affective computing benchmarks demonstrate that MFMC consistently achieves state-of-the-art or competitive performance under both subject-dependent and subject-independent evaluation protocols, highlighting its robustness to inter-subject variability. In particular, MFMC improves subject-dependent accuracy on CEAP-360VR from 78.9% to 86.8%, and subject-independent accuracy from 27.5% to 33.1% using the EDA signal alone. Moreover, MFMC remains within 0.8 percentage points of the best-performing method on the most challenging EEG subject-independent split of MAHNOB-HCI. Our code is available at https://github.com/DY9910/MFMC.

情感识别多模态学习自监督生理信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。