arXiv:2605.08252cs.CV2026-05

解决多模态情感识别中多数类主导问题,提升小样本情绪识别效果。

Multimodal Emotion Recognition via Causal-Diffusion Bridge (Affect-Diff)

  • 构建因果-扩散桥梁,动态调整模态权重并约束潜在空间结构。
  • 在CMU-MOSEI上平衡准确率达0.384,比最强基线提升18%。
  • 仅确定性编码器可识别全部6类情绪,揭示KL正则强度关键作用。

CMU-MOSEI数据集上的多模态情感识别面临极端不平衡:快乐类占65.9%,而埃克曼六类情绪合计不足7%。标准融合模型为最大化准确率而完全忽略少数类。我们提出Affect-Diff,一种联合训练的因果-扩散桥接框架,包含三个机制:通过NOTEARS学习的因果图在融合前重加权模态贡献,使用beta-VAE瓶颈实现正则化潜在压缩,以及采用停梯度1D DDPM先验来防止多数类坍缩。在3,292个对齐的CMU-MOSEI样本上,Affect-Diff验证集平衡准确率达0.384,相较最强基线TETFN(0.324)提升18%;所有对比模型在恐惧、厌恶和惊讶三类上F1均为0。消融实验表明,扩散先验(无则下降24%)与因果图(无则下降13%)均具独立非冗余贡献。值得注意的是,仅确定性编码器版本能检测全部六类情绪,揭示KL正则强度是调控少数类敏感性的直接杠杆。

原文摘要 · Abstract (English)

Multimodal emotion recognition on CMU-MOSEI faces an extreme imbalance as Happy accounts for 65.9% of samples while three Ekman categories collectively represent under 7%, causing standard fusion models to maximize accuracy by ignoring minority emotions entirely. We present Affect-Diff, a Causal-Diffusion Bridge that addresses this through three jointly trained mechanisms: a NOTEARS-learned causal graph that re-weights modality contributions before fusion, a beta-VAE bottleneck for regularized latent compression, and a stop-gradiented 1D DDPM prior that structures the latent space against majority-class collapse. On 3,292 aligned CMU-MOSEI samples, Affect-Diff achieves validation balanced accuracy 0.384, an 18% relative improvement over the strongest baseline (TETFN: 0.324), while all evaluated baselines produce zero F1 on Fear, Disgust, and Surprise. Ablation studies confirm independent, non-redundant contributions from the diffusion prior (-24% without it) and causal graph (-13%). Notably, only the deterministic-encoder variant detects all six emotion classes, revealing KL regularization strength as a direct lever for minority-class sensitivity.

情感识别多模态扩散模型因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。