用SHAP权重融合多模态情感识别专家,提升效果同时保持可解释性。
SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits

- 基于TreeSHAP计算样本级权重,动态融合单模与跨模态专家
- 在MELD和CMU-MOSEI上接近早期融合性能,优于传统晚期融合
- 揭示了不同SHAP归约方式对高维跨模态专家的影响机制
多模态情感与情绪识别通常采用早期融合(拼接特征后分类)或晚期融合(独立训练单模模型再组合)。前者准确但缺乏模块性,后者模块化但易丢失跨模态交互。本文重新审视基于XAI的自适应融合方法(—xgaf—),其样本级权重由TreeSHAP贡献度生成。研究聚焦于专家维度不均时的SHAP归约效应:均值绝对值与中位数绝对值归约会抑制高维跨模态专家,而求和绝对值归约能保持总贡献质量。在MELD 7类情绪识别任务中,求和绝对值 —xgaf— 接近早期融合表现,三种人脸序列聚合器下最高达0.5983 —wf—,仅略低于早期融合的0.6018,显著优于概率平均晚期融合的0.4598。McNemar检验显示 —xgaf— 与早期融合无显著差异(p=1.000),但优于晚期融合(p<0.0001)。在CMU-MOSEI 3类情感识别中,求和绝对值 —xgaf— 达0.6519 —wf—,略超早期融合(0.6485)和晚期融合(0.5696)。消融实验表明,性能提升主要来自引入跨模态专家(尤其是三模态专家),而非复杂的逐样本路由。诊断分析显示,均值/中位数绝对值权重趋于均匀,而求和绝对值权重集中于三模态专家。核心贡献在于透明地揭示了SHAP归约、专家维度与跨模态设计对模块化多模态融合的影响。
原文摘要 · Abstract (English)
Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors. Early fusion can be accurate but monolithic, while late fusion is modular but may lose cross-modal interactions. This paper revisits XAI-guided adaptive fusion (\xgaf), a tree-based mixture of unimodal and cross-modal experts whose sample-level weights are derived from TreeSHAP attribution magnitudes. We focus on the effect of SHAP attribution reduction when experts have unequal feature dimensionalities. In this setting, mean-abs and median-abs reductions can suppress high-dimensional cross-modal experts, whereas sum-abs reduction preserves total attribution mass. On MELD 7-class emotion recognition, sum-abs \xgaf{} nearly matches early fusion across three face-sequence aggregators; the Transformer variant reaches 0.5983 \wf{}, compared with 0.6018 for early fusion and 0.4598 for probability-average late fusion. McNemar testing shows no significant difference between sum-abs \xgaf{} and early fusion on MELD ($p=1.000$), while \xgaf{} remains significantly better than late fusion ($p<0.0001$). On CMU-MOSEI 3-class sentiment recognition, sum-abs \xgaf{} reaches 0.6519 \wf{}, slightly exceeding early fusion (0.6485) and late fusion (0.5696). Ablation studies show that the main gain comes from adding cross-modal experts, especially the trimodal expert, rather than from complex per-sample routing. Diagnostics further show that mean-abs and median-abs weights are nearly uniform, while sum-abs weights concentrate on the trimodal expert. Thus, the main contribution is a transparent empirical analysis of how SHAP reduction, expert dimensionality, and cross-modal expert design affect modular multimodal fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。