让多模态模型自动判断哪个信息更可靠,动态调整融合方式。
Learning to Fuse: Modality-Aware Adaptive Scheduling for Robust Multimodal Foundation Models
- 用视觉与文本熵、跨模态一致性信号,动态预测各模态权重。
- 在图像检索、图文生成等任务上超越CLIP、ALBEF等基线模型。
- 特别适合处理噪声、缺失或错位的输入数据,提升鲁棒性。
多模态基础模型在众多视觉-语言任务中取得显著进展,但现有方法通常采用固定或任务特定的融合策略,忽视了模态可靠性与样本复杂性的内在差异。本文提出模态感知自适应调度融合(MA-AFS),一种通用框架,可按实例动态调节各模态的贡献。该框架引入轻量级神经调度器,通过整合视觉与文本熵信号及跨模态一致性线索,预测模态融合权重。这使模型能自适应强调更可靠的模态,尤其在噪声、缺失或错位输入下表现优异。我们将融合过程建模为可微调度机制,分析其理论一致性与正则化效果,表明其在不显著增加模型容量的前提下提升鲁棒性。在图像-文本检索、图像描述生成和视觉问答任务上的大量实验显示,MA-AFS持续优于CLIP、ALBEF和BLIP等强基线。此外,它在模态损坏下表现更优,并在领域迁移时具备更强泛化能力。本工作凸显了自适应融合的重要性,为构建可靠且具备不确定性感知能力的多模态学习开辟新方向。
原文摘要 · Abstract (English)
Multimodal foundation models have achieved impressive progress across a wide range of vision-language tasks. However, existing approaches often adopt fixed or task-specific fusion strategies, neglecting the intrinsic variability of modality reliability and sample complexity. In this paper, we propose Modality-Aware Adaptive Fusion Scheduling (MA-AFS), a general framework that learns to dynamically modulate the contribution of each modality on a per-instance basis. MA-AFS introduces a lightweight neural scheduler that predicts modality fusion weights by integrating visual and textual entropy signals along with cross-modal agreement cues. This enables the model to adaptively emphasize more reliable modalities, especially under noisy, missing, or misaligned inputs. We formulate the fusion process as a differentiable scheduling mechanism, analyze its theoretical consistency and regularization effect, and demonstrate that it improves robustness without increasing model capacity significantly. Extensive experiments on image-text retrieval, captioning, and visual question answering show that MA-AFS achieves consistent performance gains over strong baselines such as CLIP, ALBEF, and BLIP. Moreover, MA-AFS exhibits improved robustness under modality corruption and enhanced generalization under domain shifts. Our work highlights the importance of adaptive fusion and opens a promising direction toward reliable and uncertainty-aware multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。