解决多模态模型在跨域泛化中因收敛速度差异导致的模态失衡问题
Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
- 通过自适应模态丢弃缓解快速收敛模态的主导倾向
- 利用梯度一致性约束实现单模态与融合表示的协同优化
- 基于权重平均的教师模型促进跨模态知识迁移,提升泛化能力
加权平均(WA)通过引导模型收敛至平坦的损失曲面,显著提升跨分布性能。然而,在多模态领域泛化(MMDG)中直接应用WA存在挑战:各模态优化速度不一,导致早期阶段过度拟合于快速收敛模态,抑制了慢速但互补模态的贡献,阻碍有效模态融合并使损失曲面趋向更尖锐、泛化性差的极小值。为此,本文提出MBCD,一种统一的协作蒸馏框架,在保留WA平滑损失曲面优势的同时,克服其在多模态场景下的缺陷。MBCD首先在学生模型中引入自适应模态丢弃,以抑制早期对主导模态的偏倚;随后通过梯度一致性约束,对齐单模态分支与融合表示的学习信号,促进协调平滑的优化过程;最后,基于WA的教师模型通过将融合知识蒸馏至各单模态分支,强化跨模态交互,引导收敛至更平坦的解空间。在多个MMDG基准上的大量实验表明,MBCD持续优于现有方法,在多种未见领域上均取得更高的准确率与鲁棒性。
原文摘要 · Abstract (English)
Weight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying WA directly to multi-modal domain generalization (MMDG) is challenging: differences in optimization speed across modalities lead WA to overfit to faster-converging ones in early stages, suppressing the contribution of slower yet complementary modalities, thereby hindering effective modality fusion and skewing the loss surface toward sharper, less generalizable minima. To address this issue, we propose MBCD, a unified collaborative distillation framework that retains WA's flatness-inducing advantages while overcoming its shortcomings in multi-modal contexts. MBCD begins with adaptive modality dropout in the student model to curb early-stage bias toward dominant modalities. A gradient consistency constraint then aligns learning signals between uni-modal branches and the fused representation, encouraging coordinated and smoother optimization. Finally, a WA-based teacher conducts cross-modal distillation by transferring fused knowledge to each uni-modal branch, which strengthens cross-modal interactions and steer convergence toward flatter solutions. Extensive experiments on MMDG benchmarks show that MBCD consistently outperforms existing methods, achieving superior accuracy and robustness across diverse unseen domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。