arXiv:2505.22209cs.CV2025-05综述

不训练模型,用现成多模态模型实现开放词汇语义分割

A Survey on Training-free Open-Vocabulary Semantic Segmentation

论文配图:A Survey on Training-free Open-Vocabulary Semantic Segmentation
图 1 · 摘自论文原文
  • 利用预训练多模态模型,无需新数据训练即可分割未知类别
  • 涵盖30+方法,按是否依赖CLIP、视觉基础模型或生成模型分类
  • 适合刚入行的研究者,提供领域全景与未来方向

语义分割是图像理解中的一项基础任务,研究历史悠久,方法众多。传统方法需从头训练模型,耗费大量计算资源和标注数据。在开放词汇语义分割场景下,要求模型识别未见过的类别,而精细标注数据成本过高。因此,研究者转向训练自由方法,利用已有的、针对易获取数据的任务训练的预训练模型。本文综述了训练自由开放词汇语义分割的发展历程、核心思想、技术演进与最新进展,重点分析基于CLIP的方法、借助辅助视觉基础模型的方法以及依赖生成模型的方法。全文系统梳理了超过30种代表性方法,按研究分支分类,并讨论当前方法的局限性与潜在问题,提出若干未充分探索的研究方向。本综述可作为新研究者的入门读物,激发对该领域的兴趣。

原文摘要 · Abstract (English)

Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional methods strive to train models up from scratch, requiring vast amounts of computational resources and training data. In the advent of moving to open-vocabulary semantic segmentation, which asks models to classify beyond learned categories, large quantities of finely annotated data would be prohibitively expensive. Researchers have instead turned to training-free methods where they leverage existing models made for tasks where data is more easily acquired. Specifically, this survey will cover the history, nuance, idea development and the state-of-the-art in training-free open-vocabulary semantic segmentation that leverages existing multi-modal classification models. We will first give a preliminary on the task definition followed by an overview of popular model archetypes and then spotlight over 30 approaches split into broader research branches: purely CLIP-based, those leveraging auxiliary visual foundation models and ones relying on generative methods. Subsequently, we will discuss the limitations and potential problems of current research, as well as provide some underexplored ideas for future study. We believe this survey will serve as a good onboarding read to new researchers and spark increased interest in the area.

语义分割开放词汇训练自由CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。