arXiv:2501.18592cs.CVcs.AI2025-01TPAMI被引 40

系统梳理多模态自适应与泛化从传统方法到大模型的演进路径

Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models

  • 按问题类型分五类梳理多模态自适应技术体系
  • 涵盖域适应、测试时自适应、泛化等核心场景
  • 适合关注多模态模型鲁棒性与迁移能力的研究者

现实场景中,模型需适应或泛化至未知目标分布,尤其在跨模态分布时挑战更大。近年来研究覆盖动作识别、语义分割等多个应用,从传统方法发展至基于CLIP等大规模预训练多模态基础模型的新范式。本文首次全面综述多模态自适应与泛化进展,包括:(1)多模态域适应;(2)多模态测试时自适应;(3)多模态域泛化;(4)利用多模态基础模型提升自适应性能;(5)多模态基础模型的适配。对每类问题进行形式化定义并系统回顾方法,分析相关数据集与应用,指出开放挑战与未来方向。持续维护文献库:https://github.com/donghao51/Awesome-Multimodal-Adaptation。

原文摘要 · Abstract (English)

In real-world scenarios, achieving domain adaptation and generalization poses significant challenges, as models must adapt to or generalize across unknown target distributions. Extending these capabilities to unseen multimodal distributions, i.e., multimodal domain adaptation and generalization, is even more challenging due to the distinct characteristics of different modalities. Significant progress has been made over the years, with applications ranging from action recognition to semantic segmentation. Besides, the recent advent of large-scale pre-trained multimodal foundation models, such as CLIP, has inspired works leveraging these models to enhance adaptation and generalization performances or adapting them to downstream tasks. This survey provides the first comprehensive review of recent advances from traditional approaches to foundation models, covering: (1) Multimodal domain adaptation; (2) Multimodal test-time adaptation; (3) Multimodal domain generalization; (4) Domain adaptation and generalization with the help of multimodal foundation models; and (5) Adaptation of multimodal foundation models. For each topic, we formally define the problem and thoroughly review existing methods. Additionally, we analyze relevant datasets and applications, highlighting open challenges and potential future research directions. We maintain an active repository that contains up-to-date literature at https://github.com/donghao51/Awesome-Multimodal-Adaptation.

多模态域适应基础模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。