arXiv:2412.09240cs.CVcs.AI2024-12被引 4

用无监督域适应提升视觉语言模型的细粒度分割能力

VLMs meet UDA: Boosting Transferability of Open Vocabulary Segmentation with Unsupervised Domain Adaptation

  • 融合多尺度上下文与提示增强,改进VLM细粒度理解
  • 通过知识蒸馏和跨域混合采样,实现无共享类别域自适应
  • 首个无需共用类别的无监督域适应分割框架,适合跨域场景

分割模型通常受限于训练时定义的类别。为解决此问题,研究者分别探索了视觉语言模型(VLMs)和合成数据两种独立路径。但VLMs在细粒度概念分离上表现不佳,而基于合成数据的方法则受限于现有数据集范围。本文提出将视觉语言推理与无监督域适应(UDA)关键策略结合,以提升跨域分割精度。首先,在提出的基础保留式开放词汇语义分割(FROVSS)框架中,通过多尺度上下文数据、鲁棒文本嵌入与提示增强,以及分层微调,增强VLM的细粒度分割能力。其次,将上述改进整合进UDA框架,采用知识蒸馏稳定训练,并通过跨域混合采样提升适应性而不牺牲泛化能力。所提出的UDA-FROVSS框架是首个可在无共享类别条件下有效跨域适配的UDA方法。

原文摘要 · Abstract (English)

Segmentation models are typically constrained by the categories defined during training. To address this, researchers have explored two independent approaches: adapting Vision-Language Models (VLMs) and leveraging synthetic data. However, VLMs often struggle with granularity, failing to disentangle fine-grained concepts, while synthetic data-based methods remain limited by the scope of available datasets. This paper proposes enhancing segmentation accuracy across diverse domains by integrating Vision-Language reasoning with key strategies for Unsupervised Domain Adaptation (UDA). First, we improve the fine-grained segmentation capabilities of VLMs through multi-scale contextual data, robust text embeddings with prompt augmentation, and layer-wise fine-tuning in our proposed Foundational-Retaining Open Vocabulary Semantic Segmentation (FROVSS) framework. Next, we incorporate these enhancements into a UDA framework by employing distillation to stabilize training and cross-domain mixed sampling to boost adaptability without compromising generalization. The resulting UDA-FROVSS framework is the first UDA approach to effectively adapt across domains without requiring shared categories.

语义分割视觉语言模型无监督域适应开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。