arXiv:2601.20064cs.CV2026-01被引 1

提出DiSa框架,解决开放词汇语义分割中背景忽略与边界模糊问题。

DiSa: Saliency-Aware Foreground-Background Disentangled Framework for Open-Vocabulary Semantic Segmentation

  • 分治建模前景与背景特征,引入显著性感知模块增强区分能力
  • 在6个基准上超越现有方法,显著提升背景区域识别准确率
  • 适合需要精准分割非显著区域的视觉任务,如遥感与医学图像分析

开放词汇语义分割旨在根据文本标签为图像中每个像素分配类别。现有方法通常利用视觉语言模型(如CLIP)进行密集预测,但这些模型在图像-文本对上预训练,偏向显著的、以物体为中心的区域,在适应分割任务时存在两大缺陷:(i) 前景偏差,忽略背景区域;(ii) 空间定位能力有限,导致物体边界模糊。为此,我们提出DiSa,一种新颖的显著性感知前景-背景解耦框架。通过设计的显著性感知解耦模块(SDM),DiSa显式引入显著性线索,以分治方式分别建模前景与背景的集成特征。此外,我们提出层次化精修模块(HRM),利用像素级空间上下文,通过多层级更新实现通道级特征精修。在六个基准上的大量实验表明,DiSa持续优于当前最优方法。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation aims to assign labels to every pixel in an image based on text labels. Existing approaches typically utilize vision-language models (VLMs), such as CLIP, for dense prediction. However, VLMs, pre-trained on image-text pairs, are biased toward salient, object-centric regions and exhibit two critical limitations when adapted to segmentation: (i) Foreground Bias, which tends to ignore background regions, and (ii) Limited Spatial Localization, resulting in blurred object boundaries. To address these limitations, we introduce DiSa, a novel saliency-aware foreground-background disentangled framework. By explicitly incorporating saliency cues in our designed Saliency-aware Disentanglement Module (SDM), DiSa separately models foreground and background ensemble features in a divide-and-conquer manner. Additionally, we propose a Hierarchical Refinement Module (HRM) that leverages pixel-wise spatial contexts and enables channel-wise feature refinement through multi-level updates. Extensive experiments on six benchmarks demonstrate that DiSa consistently outperforms state-of-the-art methods.

语义分割视觉语言模型前景背景解耦显著性感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。