arXiv:2605.28348cs.CV2026-05中稿 · the 2026 IEEE Inte…

让视觉语言分割模型摆脱语义依赖,学会凭形状纹理等视觉特征分割

Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models

论文配图:Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
图 1 · 摘自论文原文
  • 用非语义文本描述生成提示,训练模型忽略类别信息
  • 在新任务上比现有模型提升20%分割准确率
  • 适合需要精细控制和通用分割的场景

视觉语言分割模型近期通过利用自然语言表达的高层语义类别取得了优异性能。然而,这种对语义的依赖限制了其对形状、几何或纹理等内在视觉属性的推理能力,而这些属性在许多实际应用中至关重要。本文提出语义无偏且形状感知(SANSA)分割新范式,要求分割模型仅基于非语义文本描述进行操作。为此,我们提出两种生成SANSA分割提示的方法:基于词典约束或示例引导,均生成语义无偏的文本描述。这些提示用于在语义无偏监督下微调分割模型。实验表明,在新任务上,使用SANSA提示微调可使模型的平均交并比(mIoU)相比预训练先进模型提升最高达20%,同时保持在标准语义提示上的良好性能。结果凸显了低级和中级视觉推理对提升视觉语言分割模型泛化性和可控性的关键作用。

原文摘要 · Abstract (English)

Vision-language segmentation models have recently achieved strong performance by leveraging high-level semantic object categories expressed in natural language. However, this semantic dependence limits their ability to reason about intrinsic visual properties such as shape, geometry, or texture, which are essential in many real-world applications. In this work, we introduce Semantic-Agnostic aNd Shape-Aware (SANSA) segmentation, a new paradigm that requires segmentation models to operate solely from non-semantic textual descriptions. To this end, we propose two strategies to generate SANSA segmentation prompts based on either dictionary constraints or example guidance, both generating semantic-agnostic textual descriptions. These prompts are then used to finetune segmentation models under semantic-agnostic supervision. Experiments show that finetuning on SANSA prompts yields up to a 20% mIoU improvement on this new segmentation task, compared to pretrained state-of-the-art models, while maintaining strong performance on standard semantic prompts. These results highlight the importance of low- and mid-level visual reasoning for improving the generalization and controllability of vision-language segmentation models.

视觉分割多模态语义无偏形状感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。