arXiv:2508.05547cs.LGcs.AI2025-08IJCV综述被引 10

不依赖标签,四类无监督方法让视觉语言模型更好适应新任务。

Adapting Vision-Language Models Without Labels: A Comprehensive Survey

  • 按无标签数据形式分四类:无数据、大量数据、批量数据、流式数据
  • 提出系统框架,梳理各方法的核心策略与适用场景
  • 适合关注模型泛化与数据效率的研究者和工程师

视觉语言模型在多种任务中表现出色,但在未经过特定任务适配的情况下性能常不理想。为提升其实用性并保持数据高效性,近期研究聚焦于无需标注数据的无监督适配方法。尽管该领域关注度上升,仍缺乏面向任务的统一综述。为此,本文提出一个基于无标签视觉数据可用性与性质的分类体系,将现有方法分为四类:无数据迁移(Data-Free Transfer)、无监督域迁移(Unsupervised Domain Transfer)、短时测试自适应(Episodic Test-Time Adaptation)和在线测试自适应(Online Test-Time Adaptation)。在此框架下,分析各类方法的核心机制与适配策略,建立系统性理解。同时,综述多个代表性基准及应用场景,指出当前挑战与未来方向。相关文献库持续维护于 https://github.com/tim-learn/Awesome-LabelFree-VLMs。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable generalization capabilities across a wide range of tasks. However, their performance often remains suboptimal when directly applied to specific downstream scenarios without task-specific adaptation. To enhance their utility while preserving data efficiency, recent research has increasingly focused on unsupervised adaptation methods that do not rely on labeled data. Despite the growing interest in this area, there remains a lack of a unified, task-oriented survey dedicated to unsupervised VLM adaptation. To bridge this gap, we present a comprehensive and structured overview of the field. We propose a taxonomy based on the availability and nature of unlabeled visual data, categorizing existing approaches into four key paradigms: Data-Free Transfer (no data), Unsupervised Domain Transfer (abundant data), Episodic Test-Time Adaptation (batch data), and Online Test-Time Adaptation (streaming data). Within this framework, we analyze core methodologies and adaptation strategies associated with each paradigm, aiming to establish a systematic understanding of the field. Additionally, we review representative benchmarks across diverse applications and highlight open challenges and promising directions for future research. An actively maintained repository of relevant literature is available at https://github.com/tim-learn/Awesome-LabelFree-VLMs.

视觉语言模型无监督学习模型适配数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。