arXiv:2507.03304cs.CV2025-07ICCV被引 10

用统一表示让多模态模型在未知场景下更稳定,提升跨域泛化能力。

Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

  • 通过统一表示空间同步优化多模态特征,避免各模态方向偏离。
  • 在EPIC-Kitchens和Human-Animal-Cartoon上达到领先性能,显著提升泛化效果。
  • 适合需要跨域鲁棒性的多模态任务,如视频理解、人机交互等。

领域泛化(DG)旨在仅通过源域训练,提升模型在未见或分布偏移目标域中的鲁棒性。尽管现有方法在单模态数据上取得进展,但多数难以直接应用于多模态场景。多模态领域泛化(MMDG)面临核心挑战:如何使基于多模态源域训练的模型,在相同模态组合下泛化至未见目标分布。由于模态间差异,直接迁移单模态方法常导致结果次优,且因目标域不可见而出现随机性。独立处理各模态再融合,易引发不同模态间泛化方向不一致,削弱整体性能。为此,本文提出一种新方法,利用统一表示将多模态配对映射至共享空间,实现跨模态协同优化,有效适配经典DG方法至MMDG。同时引入监督解耦框架,分离模态共性与特异性信息,增强统一表示的一致性。在EPIC-Kitchens与Human-Animal-Cartoon等基准数据集上的实验表明,该方法显著优于现有方法,具备更强的多模态领域泛化能力。

原文摘要 · Abstract (English)

Domain Generalization (DG) aims to enhance model robustness in unseen or distributionally shifted target domains through training exclusively on source domains. Although existing DG techniques, such as data manipulation, learning strategies, and representation learning, have shown significant progress, they predominantly address single-modal data. With the emergence of numerous multi-modal datasets and increasing demand for multi-modal tasks, a key challenge in Multi-modal Domain Generalization (MMDG) has emerged: enabling models trained on multi-modal sources to generalize to unseen target distributions within the same modality set. Due to the inherent differences between modalities, directly transferring methods from single-modal DG to MMDG typically yields sub-optimal results. These methods often exhibit randomness during generalization due to the invisibility of target domains and fail to consider inter-modal consistency. Applying these methods independently to each modality in the MMDG setting before combining them can lead to divergent generalization directions across different modalities, resulting in degraded generalization capabilities. To address these challenges, we propose a novel approach that leverages Unified Representations to map different paired modalities together, effectively adapting DG methods to MMDG by enabling synchronized multi-modal improvements within the unified space. Additionally, we introduce a supervised disentanglement framework that separates modal-general and modal-specific information, further enhancing the alignment of unified representations. Extensive experiments on benchmark datasets, including EPIC-Kitchens and Human-Animal-Cartoon, demonstrate the effectiveness and superiority of our method in enhancing multi-modal domain generalization.

多模态领域泛化统一表示跨域鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。