让分割模型理解用户自定义概念,如'我的马克杯'
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
- 用少量图像和掩码,通过文本提示调优识别个性化视觉概念
- 引入负掩码提议机制,有效降低误检率
- 适合需要个性化图像分割的用户场景
开放词汇语义分割(OVSS)虽能基于任意文本描述分割图像中的语义区域,包括训练中未见的类别,但难以理解用户自定义的文本(如‘我的马克杯’)来定位特定兴趣区域。本文针对此类问题,提出个性化开放词汇语义分割新任务,并设计一种基于文本提示调优的插件方法,仅需少量图像-掩码对即可识别个性化视觉概念,同时保持原有OVSS性能。为减少提示调优带来的误检,方法引入‘负掩码提议’以捕捉非目标概念;并通过将个性化概念的视觉嵌入注入文本提示,增强语义表达。实验在新构建的FSS^per、CUB^per、ADE^per基准上验证了方法优越性。
原文摘要 · Abstract (English)
While open-vocabulary semantic segmentation (OVSS) can segment an image into semantic regions based on arbitrarily given text descriptions even for classes unseen during training, it fails to understand personal texts (e.g., `my mug cup') for segmenting regions of specific interest to users. This paper addresses challenges like recognizing `my mug cup' among `multiple mug cups'. To overcome this challenge, we introduce a novel task termed \textit{personalized open-vocabulary semantic segmentation} and propose a text prompt tuning-based plug-in method designed to recognize personal visual concepts using a few pairs of images and masks, while maintaining the performance of the original OVSS. Based on the observation that reducing false predictions is essential when applying text prompt tuning to this task, our proposed method employs `negative mask proposal' that captures visual concepts other than the personalized concept. We further improve the performance by enriching the representation of text prompts by injecting visual embeddings of the personal concept into them. This approach enhances personalized OVSS without compromising the original OVSS performance. We demonstrate the superiority of our method on our newly established benchmarks for this task, including FSS$^\text{per}$, CUB$^\text{per}$, and ADE$^\text{per}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。