arXiv:2608.26879cs.LGcs.MM2026-08

解决多模态融合中强主导模态性能下降问题,提升融合效果。

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

  • 采用反向非对称融合,让主导模态不受干扰,弱模态向其对齐
  • 在MultiHuSE上使主导模态准确率保持74.9%,避免18.5%的性能下降
  • 适用于文本主导或音视频主导的多模态任务,尤其适合强模态易被压制场景

多模态融合本应提升模型表现,但在MultiHuSE数据集上,早期、晚期及对称注意力融合均未能超越最优单模态基线(文本)。路径隔离分析显示,对称注意力融合后文本路径准确率从74.9%降至56.4%,表明主导模态在融合过程中被削弱。我们称此为强模态坍塌,并认为这是部分多模态模型无法超越单模态基线的原因。为此提出反向非对称融合(IAF),避免模态间强制相互注意力。主导模态在融合中保持不变,弱模态则以它为上下文锚点进行关注。融合前通过模态感知知识蒸馏增强弱模态。在三个具有不同模态层级的数据集(MultiHuSE、UR-FUNNY、MUStARD)上评估,路径隔离结果显示,IAF在所有配置下均将主导模态内部准确率维持在单模态上限,而对称融合在MultiHuSE上使其下降最多达18.5%。IAF相较最强单模态基线提升最高达8.25%。

原文摘要 · Abstract (English)

Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9% to 56.4% after fusion in one such setting, indicating that the dominant modality can be degraded during integration. We term this strong-modality collapse and argue that it helps explain why some multimodal models fail to surpass unimodal baselines. We propose Inverted Asymmetric Fusion (IAF), which avoids forcing mutual attention across modalities. The dominant modality is preserved by passing through fusion unchanged, while weaker modalities attend to it as a contextual anchor. Before fusion, weaker modalities are strengthened using Modality-Aware Knowledge Distillation. We evaluate IAF on three benchmarks with different modality hierarchies: text-dominant datasets (MultiHuSE, UR-FUNNY) and an audio-visual-dominant dataset (MUStARD). Pathway isolation shows that IAF preserves the dominant modality's internal accuracy at its unimodal ceiling across all tested configurations, whereas symmetric fusion degrades it by up to 18.5% on MultiHuSE. IAF improves over the strongest unimodal baseline by up to 8.25%.

多模态融合模态坍塌知识蒸馏注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。