arXiv:2512.01701cs.CV2025-12AAAI被引 3

提出SSR方法,解决CLIP在弱监督分割中的误激活问题。

SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised Segmentation

  • 通过跨模态原型对齐,优化语义特征空间分布。
  • 利用超像素引导校正,显著降低背景误激活率。
  • 在PASCAL VOC和MS COCO上分别达79.5%和50.6% mIoU。

近年来,对比语言-图像预训练(CLIP)因其强大的跨模态语义理解能力,被广泛应用于弱监督语义分割(WSSS)任务。本文提出一种新的语义与空间校正(SSR)方法,以解决现有基于CLIP的弱监督分割方法存在的问题:非目标前景区域和背景区域的过激活现象。具体而言,在语义层面,跨模态原型对齐(CMPA)建立对比学习机制,实现模态间特征空间对齐,减少类间混淆,增强语义相关性,有效校正非目标前景区域的过激活;在空间层面,超像素引导校正(SGC)利用基于超像素的空间先验,在亲和传播过程中精确过滤非目标区域干扰,显著缓解背景过激活问题。在PASCAL VOC和MS COCO数据集上的大量实验表明,该方法优于所有单阶段方法以及更复杂的多阶段方法,在两个数据集上分别取得79.5%和50.6%的mIoU性能。

原文摘要 · Abstract (English)

In recent years, Contrastive Language-Image Pretraining (CLIP) has been widely applied to Weakly Supervised Semantic Segmentation (WSSS) tasks due to its powerful cross-modal semantic understanding capabilities. This paper proposes a novel Semantic and Spatial Rectification (SSR) method to address the limitations of existing CLIP-based weakly supervised semantic segmentation approaches: over-activation in non-target foreground regions and background areas. Specifically, at the semantic level, the Cross-Modal Prototype Alignment (CMPA) establishes a contrastive learning mechanism to enforce feature space alignment across modalities, reducing inter-class overlap while enhancing semantic correlations, to rectify over-activation in non-target foreground regions effectively; at the spatial level, the Superpixel-Guided Correction (SGC) leverages superpixel-based spatial priors to precisely filter out interference from non-target regions during affinity propagation, significantly rectifying background over-activation. Extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our method outperforms all single-stage approaches, as well as more complex multi-stage approaches, achieving mIoU scores of 79.5% and 50.6%, respectively.

弱监督分割CLIP语义校正超像素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。