arXiv:2608.11681cs.CVcs.AI2026-08

用多模态伪标签提升开放词汇实例与全景分割的泛化能力

Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation

论文配图:Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation
图 1 · 摘自论文原文
  • 基于CLIP和LLaVA自动生成伪掩码与描述,避免人工标注
  • 引入同义词过滤与生成式重构,提升视觉-文本对齐精度
  • 在COCO数据集上显著优于现有方法,适合开放场景分割任务

本文针对开放词汇实例分割(OVIS)与开放集全景分割(OSPS)的挑战,旨在识别预定义及未见物体类别而无需全量人工标注。现有方法常受限于噪声伪掩码、视觉-文本对齐不足及同义词或词汇外(OOV)词处理困难。为此,提出一种多模态框架,利用预训练视觉语言模型实现自动伪标签生成、CLIP引导的同义词过滤以及GPT驱动的描述重建。在目标词汇辅助的伪标签设定下,框架通过Grounded SAM、LLaVA和CLIP构建伪分割掩码、描述性文字与语义对齐的同义词集,提供无标注的多模态监督。进一步通过三种互补训练目标增强视觉-文本对齐:扩展的接地损失(融合视觉锚定的同义词)、语义一致性损失与生成式描述重构损失。在COCO数据集上的大量实验表明,所提方法在该协议下持续超越现有最先进方法,在OVIS与OSPS基准上均取得显著提升。

原文摘要 · Abstract (English)

This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, we propose a multimodal framework that leverages pre-trained vision-language models for automatic pseudo-label generation, CLIP-guided synonym filtering, and GPT-based caption reconstruction. In our target-vocabulary-assisted pseudo-labeling setting, the framework first constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP, providing multimodal supervision without manual annotation. We then enhance visual-textual alignment through three complementary training objectives: an extended grounding loss that incorporates visually grounded synonyms, a semantic consistency loss, and a generative caption reconstruction loss. Extensive experiments on the COCO dataset demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches under this protocol, achieving substantial improvements on both OVIS and OSPS benchmarks.

实例分割全景分割多模态伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。