用视觉原型解决图文对齐难题,提升弱监督语义分割精度
Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
- 引入视觉原型学习,让图像特征与文本更好匹配
- 在两个基准数据集上达到当前最优性能
- 适合研究跨模态对齐与弱监督分割的学者
对比语言-图像预训练(CLIP)在弱监督语义分割(WSSS)中展现出强大的跨模态语义理解能力。现有方法通过微调文本提示来优化图像与文本的对齐,但受限于文本与视觉空间间的模态鸿沟,这些方法中的文本原型未能有效对应像素级视觉特征。本文理论分析表明,该模态鸿沟导致图像区域特征与文本特征错位,且仅最小化CLIP对比损失无法充分缓解。为此,我们提出视觉原型学习(VPL)框架,通过引入更具代表性的视觉原型,结合文本原型在视觉空间中学习类别特定的视觉原型,以捕捉高质量定位图。此外,设计区域语义对比模块,将区域嵌入与对应原型进行对比,实现更全面、鲁棒的特征学习。实验结果表明,所提框架在两个基准数据集上均达到领先性能。
原文摘要 · Abstract (English)
The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research powerful cross-modal semantic understanding capabilities. Existing methods attempt to optimize input text prompts for improved alignment of images and text, by finely adjusting text prototypes to facilitate semantic matching. Nevertheless, given the modality gap between text and vision spaces, the text prototypes employed by these methods have not effectively established a close correspondence with pixel-level vision features. In this work, our theoretical analysis indicates that the inherent modality gap results in misalignment of text and region features, and that this gap cannot be sufficiently reduced by minimizing contrast loss in CLIP. To mitigate the impact of the modality gap, we propose a Vision Prototype Learning (VPL) framework, by introducing more representative vision prototypes. The core of this framework is to learn class-specific vision prototypes in vision space with the help of text prototypes, for capturing high-quality localization maps. Moreover, we propose a regional semantic contrast module that contrasts regions embedding with corresponding prototypes, leading to more comprehensive and robust feature learning. Experimental results show that our proposed framework achieves state-of-the-art performance on two benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。