arXiv:2603.14175cs.LGcs.CV2026-03AAAI被引 2

提出GMP方法解决多模态模型训练不均衡问题

Balancing Multimodal Domain Generalization via Gradient Modulation and Projection

  • 分离分类与域不变性梯度,按语义和域置信度动态调节
  • 在多个基准上实现当前最优跨域泛化性能
  • 适合需要提升多模态模型鲁棒性的研究者使用

多模态领域泛化(MMDG)利用多模态间的互补优势提升模型在未见域上的泛化能力。其核心挑战是优化不平衡:不同模态在训练中收敛速度不一,导致梯度贡献不均,部分模态主导学习过程而其他模态滞后。现有平衡策略通常依据源域分类表现调节各模态梯度,但忽视了关键洞察:在源域表现好的模态可能在未见域泛化差,限制跨域收益。为此,本文提出梯度调制投影(GMP),一种统一的平衡策略。GMP首先解耦分类与域不变性目标的梯度,再基于语义和域置信度调节各模态梯度;同时通过追踪任务相对强度动态调整梯度投影,缓解模态特异性编码器内分类与域不变学习间的冲突。大量实验表明,GMP达到当前最优性能,并可灵活集成于多种MMDG方法,在多个基准上显著提升跨域泛化能力。

原文摘要 · Abstract (English)

Multimodal Domain Generalization (MMDG) leverages the complementary strengths of multiple modalities to enhance model generalization on unseen domains. A central challenge in multimodal learning is optimization imbalance, where modalities converge at different speeds during training. This imbalance leads to unequal gradient contributions, allowing some modalities to dominate the learning process while others lag behind. Existing balancing strategies typically regulate each modality's gradient contribution based on its classification performance on the source domain to alleviate this issue. However, relying solely on source-domain accuracy neglects a key insight in MMDG: modalities that excel on the source domain may generalize poorly to unseen domains, limiting cross-domain gains. To overcome this limitation, we propose Gradient Modulation Projection (GMP), a unified strategy that promotes balanced optimization in MMDG. GMP first decouples gradients associated with classification and domain-invariance objectives. It then modulates each modality's gradient based on semantic and domain confidence. Moreover, GMP dynamically adjusts gradient projections by tracking the relative strength of each task, mitigating conflicts between classification and domain-invariant learning within modality-specific encoders. Extensive experiments demonstrate that GMP achieves state-of-the-art performance and integrates flexibly with diverse MMDG methods, significantly improving generalization across multiple benchmarks.

多模态域泛化梯度调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。