用少量标注图像提升开放词汇语义分割精度
Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- 用支持集图像增强文本提示,实现图文特征动态融合
- 在多个数据集上接近监督学习性能,零样本到有监督差距缩小60%以上
- 适合个性化分割、细粒度识别等需灵活适配场景
开放词汇语义分割(OVS)将视觉语言模型(VLM)的零样本识别能力扩展至像素级预测,可对任意文本提示类别进行分割。尽管已有进展,但其性能仍落后于全监督方法,主要受限于训练时使用粗粒度图像级监督以及自然语言的语义模糊性。本文提出一种少样本设置,通过引入带有像素标注的支持图像集来增强文本提示。在此基础上,提出一种检索增强的测试时适配器,通过融合文本与视觉支持特征,为每张图像学习一个轻量级分类器。与以往依赖人工设计后期融合的方法不同,本方法采用可学习的查询级融合,实现模态间更强协同。该方法支持持续扩展支持集,适用于细粒度任务如个性化分割。实验表明,显著缩小了零样本与监督分割之间的差距,同时保持开放词汇能力。
原文摘要 · Abstract (English)
Open-vocabulary segmentation (OVS) extends the zero-shot recognition capabilities of vision-language models (VLMs) to pixel-level prediction, enabling segmentation of arbitrary categories specified by text prompts. Despite recent progress, OVS lags behind fully supervised approaches due to two challenges: the coarse image-level supervision used to train VLMs and the semantic ambiguity of natural language. We address these limitations by introducing a few-shot setting that augments textual prompts with a support set of pixel-annotated images. Building on this, we propose a retrieval-augmented test-time adapter that learns a lightweight, per-image classifier by fusing textual and visual support features. Unlike prior methods relying on late, hand-crafted fusion, our approach performs learned, per-query fusion, achieving stronger synergy between modalities. The method supports continually expanding support sets, and applies to fine-grained tasks such as personalized segmentation. Experiments show that we significantly narrow the gap between zero-shot and supervised segmentation while preserving open-vocabulary ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。