arXiv:2501.18753cs.CV2025-01IJCAI被引 5

通过自适应筛选负样本,提升通用提示的图像分割精度。

INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation

  • 基于负样本挖掘,动态过滤无关先验信息以优化实例提示生成。
  • 在6个数据集上实现更准确的分割,尤其在伪装物体和医学图像中表现优异。
  • 适合需要高鲁棒性、通用提示分割的科研与医疗场景。

任务通用的可提示图像分割旨在仅用一个通用任务提示,实现对多种样本的分割。现有方法利用视觉语言模型(VLM)从通用提示中推断出实例特定提示以指导分割过程。然而当VLM难以泛化到某些图像实例时,生成的实例提示质量较差。为此,本文提出一种新方法INT(Instance-Specific Negative Mining for Task-Generic Promptable Segmentation),其核心思想是自适应减少无关(负)先验知识的影响,同时增强通过高对比度负样本挖掘选出的最合理先验知识,以优化实例提示生成。INT包含两个组件:(1) 实例特定提示生成,逐步过滤提示生成中的错误信息;(2) 语义掩码生成,确保每个图像实例的分割结果正确匹配实例提示的语义。在包括伪装物体和医学图像在内的六个数据集上验证了INT的有效性、鲁棒性和可扩展性。

原文摘要 · Abstract (English)

Task-generic promptable image segmentation aims to achieve segmentation of diverse samples under a single task description by utilizing only one task-generic prompt. Current methods leverage the generalization capabilities of Vision-Language Models (VLMs) to infer instance-specific prompts from these task-generic prompts in order to guide the segmentation process. However, when VLMs struggle to generalise to some image instances, predicting instance-specific prompts becomes poor. To solve this problem, we introduce \textbf{I}nstance-specific \textbf{N}egative Mining for \textbf{T}ask-Generic Promptable Segmentation (\textbf{INT}). The key idea of INT is to adaptively reduce the influence of irrelevant (negative) prior knowledge whilst to increase the use the most plausible prior knowledge, selected by negative mining with higher contrast, in order to optimise instance-specific prompts generation. Specifically, INT consists of two components: (1) instance-specific prompt generation, which progressively fliters out incorrect information in prompt generation; (2) semantic mask generation, which ensures each image instance segmentation matches correctly the semantics of the instance-specific prompts. INT is validated on six datasets, including camouflaged objects and medical images, demonstrating its effectiveness, robustness and scalability.

图像分割提示学习负样本挖掘视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。