提出可插拔模块,解决多模态模型缺模态时性能崩溃问题
Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
- 通过频域分析量化模态偏好,动态调整各模态贡献权重
- 在多个模型架构和任务上实现稳定性能提升,最高增益达7.2%
- 适合需要鲁棒多模态理解的场景,如医疗影像与文本融合
多模态模型在缺失模态时常出现灾难性性能下降。我们发现这种脆弱性源于不平衡的学习过程,模型对某些模态产生隐式偏好,导致其他模态优化不足。为此,我们提出一种简单高效的方法:首先引入频域特征分析的频率比度量(FRM),量化模态偏好;再设计一个可插拔的多模态权重分配模块(MWAM),在训练中动态重平衡各分支贡献,促进更全面的学习。大量实验表明,MWAM可无缝集成至CNN与ViT等主流架构,在多种任务和模态组合中保持一致增益,不仅提升基础模型性能,还能进一步增强现有缺模态解决方案的效果。
原文摘要 · Abstract (English)
Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the under-optimization of others. We propose a simple yet efficient method to address this challenge. The central insight of our work is that the dominance relationship between modalities can be effectively discerned and quantified in the frequency domain. To leverage this principle, we first introduce a Frequency Ratio Metric (FRM) to quantify modality preference by analyzing features in the frequency domain. Guided by FRM, we then propose a Multimodal Weight Allocation Module, a plug-and-play component that dynamically re-balances the contribution of each branch during training, promoting a more holistic learning paradigm. Extensive experiments demonstrate that MWAM can be seamlessly integrated into diverse architectural backbones, such as those based on CNNs and ViTs. Furthermore, MWAM delivers consistent performance gains across a wide range of tasks and modality combinations. This advancement extends beyond merely optimizing the performance of the base model; it also manifests as further performance improvements to state-of-the-art methods addressing the missing modality problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。