让表现更强的模态主导多模态学习,提升整体性能
PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization
- 根据单模态表现排序,优先优化强模态
- 在多个数据集上显著优于传统平衡学习方法
- 无需改动模型结构,适合实际部署场景
多模态学习因实用性受到关注,但常因优化不足导致性能不如单模态模型。现有方法归因于模态间学习不平衡,通过梯度调制解决。本文认为,平衡学习并非最优,由表现更优的单模态驱动的不平衡学习反而能提升多模态性能。问题根源在于对优势模态的训练不足。为此提出性能主导模态优先(PDMP)策略:先通过独立训练的单模态模型表现排序识别优势模态,再引入非对称系数调节各模态梯度,使优势模态主导优化过程。由于仅依赖单模态表现排序,该方法与多模态模型结构和融合方式无关,具备良好实用性。大量实验验证了其有效性。
原文摘要 · Abstract (English)
Multimodal learning has attracted increasing attention due to its practicality. However, it often suffers from insufficient optimization, where the multimodal model underperforms even compared to its unimodal counterparts. Existing methods attribute this problem to the imbalanced learning between modalities and solve it by gradient modulation. This paper argues that balanced learning is not the optimal setting for multimodal learning. On the contrary, imbalanced learning driven by the performance-dominant modality that has superior unimodal performance can contribute to better multimodal performance. And the under-optimization problem is caused by insufficient learning of the performance-dominant modality. To this end, we propose the Performance-Dominant Modality Prioritization (PDMP) strategy to assist multimodal learning. Specifically, PDMP firstly mines the performance-dominant modality via the performance ranking of the independently trained unimodal model. Then PDMP introduces asymmetric coefficients to modulate the gradients of each modality, enabling the performance-dominant modality to dominate the optimization. Since PDMP only relies on the unimodal performance ranking, it is independent of the structures and fusion methods of the multimodal model and has great potential for practical scenarios. Finally, extensive experiments on various datasets validate the superiority of PDMP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。