arXiv:2506.18504cs.CVcs.AI2025-06综述被引 9

系统梳理视觉语言模型跨领域泛化方法与评测体系

Generalizing vision-language models to novel domains: A comprehensive survey

  • 按提示、参数、特征三类模块划分知识迁移方法
  • 对比主流基准上多类方法的泛化性能差异
  • 适合关注多模态模型实用落地的研究者阅读

视觉语言预训练作为融合视觉与文本模态优势的变革性技术,已发展出具备强零样本能力的视觉语言模型(VLMs)。然而在面对特定领域或专用任务时,其性能常显著下降。为此,大量研究致力于将VLM中蕴含的丰富知识迁移到下游应用中。本综述全面总结了VLM泛化相关的设置、方法、评测基准与结果。基于典型VLM结构,将现有文献分为提示式、参数式和特征式三类迁移方法,并通过回顾典型迁移学习场景,对VLM时代的迁移学习提出新解读。介绍了主流泛化评测基准,并对所涉方法进行详尽性能比较。结合大规模可泛化预训练进展,还讨论了VLM与最新多模态大语言模型(如DeepSeek-VL)的关系与差异。本文从新颖且实用的泛化视角系统梳理视觉语言研究前沿,为当前及未来多模态研究提供清晰图景。

原文摘要 · Abstract (English)

Recently, vision-language pretraining has emerged as a transformative technique that integrates the strengths of both visual and textual modalities, resulting in powerful vision-language models (VLMs). Leveraging web-scale pretraining data, these models exhibit strong zero-shot capabilities. However, their performance often deteriorates when confronted with domain-specific or specialized generalization tasks. To address this, a growing body of research focuses on transferring or generalizing the rich knowledge embedded in VLMs to various downstream applications. This survey aims to comprehensively summarize the generalization settings, methodologies, benchmarking and results in VLM literatures. Delving into the typical VLM structures, current literatures are categorized into prompt-based, parameter-based and feature-based methods according to the transferred modules. The differences and characteristics in each category are furthered summarized and discussed by revisiting the typical transfer learning (TL) settings, providing novel interpretations for TL in the era of VLMs. Popular benchmarks for VLM generalization are further introduced with thorough performance comparisons among the reviewed methods. Following the advances in large-scale generalizable pretraining, this survey also discusses the relations and differences between VLMs and up-to-date multimodal large language models (MLLM), e.g., DeepSeek-VL. By systematically reviewing the surging literatures in vision-language research from a novel and practical generalization prospective, this survey contributes to a clear landscape of current and future multimodal researches.

视觉语言模型泛化综述多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。