arXiv:2411.17150cs.CV2024-11CVPR被引 33

让分割模型理解物体上下文,精准识别未知类别

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation

  • 从视觉基础模型蒸馏谱特征,增强图像内物体一致性
  • 在多个数据集上达到当前最优性能,泛化能力强
  • 适合需要处理任意类别查询的开放词汇分割场景

开放词汇语义分割(OVSS)借助视觉语言模型(VLMs)取得进展,可通过不同学习方案对预定义类别之外的物体进行分割。尤其是无训练方法提供了可扩展、易部署的解决方案,适用于处理未见数据,是OVSS的关键目标。然而,现有方法在基于任意查询提示进行复杂物体分割时,缺乏对物体级别上下文的考虑,导致模型难以将语义一致的组件归为同一物体,并精确映射到用户定义的类别。本文提出新方法,在图像中引入物体级别上下文知识。具体而言,模型通过将视觉基础模型中的谱驱动特征蒸馏至视觉编码器的注意力机制,提升物体内部的一致性,使语义连贯的部分形成单一物体掩码。同时,利用零样本物体存在概率对文本嵌入进行优化,确保与图像中特定物体的准确对齐。通过引入物体级别上下文知识,本方法在多个数据集上达到当前最优表现,具备强大的泛化能力。

原文摘要 · Abstract (English)

Open-Vocabulary Semantic Segmentation (OVSS) has advanced with recent vision-language models (VLMs), enabling segmentation beyond predefined categories through various learning schemes. Notably, training-free methods offer scalable, easily deployable solutions for handling unseen data, a key goal of OVSS. Yet, a critical issue persists: lack of object-level context consideration when segmenting complex objects in the challenging environment of OVSS based on arbitrary query prompts. This oversight limits models' ability to group semantically consistent elements within object and map them precisely to user-defined arbitrary classes. In this work, we introduce a novel approach that overcomes this limitation by incorporating object-level contextual knowledge within images. Specifically, our model enhances intra-object consistency by distilling spectral-driven features from vision foundation models into the attention mechanism of the visual encoder, enabling semantically coherent components to form a single object mask. Additionally, we refine the text embeddings with zero-shot object presence likelihood to ensure accurate alignment with the specific objects represented in the images. By leveraging object-level contextual knowledge, our proposed approach achieves state-of-the-art performance with strong generalizability across diverse datasets.

开放词汇分割物体上下文视觉语言模型语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。