arXiv:2506.21835cs.CV2025-06ICCV被引 3

用概率提示提升SAM在视觉参考分割中的稳定性

ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic Prompts

  • 设计变分提示编码器,预测多变量提示分布
  • 在Pascal-5$^i$和COCO-20$^i$上超越现有方法
  • 适合需要鲁棒性分割的视觉任务研究者

大型基础模型的进展推动了开放集图像分割的发展,该任务旨在分割超出预定义类别的对象。在各类提示(如点、框、文本和视觉参考)中,视觉参考分割因其灵活性和强大的零样本能力脱颖而出。近期,一些基于SAM的方法通过自动生成提示引导SAM取得了显著进展。然而,这些方法常因提示编码器不完善,在目标区域边界生成提示,导致结果不稳定、鲁棒性下降。本文提出ProSAM,一种简单而有效的方法,通过学习变分提示编码器以预测多变量提示分布,避免在不稳定的区域生成提示,从而克服由低鲁棒性提示带来的问题。实验表明,ProSAM在Pascal-5$^i$和COCO-20$^i$数据集上持续优于当前最优方法,为视觉参考分割提供了更稳健的解决方案。

原文摘要 · Abstract (English)

The recent advancements in large foundation models have driven the success of open-set image segmentation, a task focused on segmenting objects beyond predefined categories. Among various prompt types (such as points, boxes, texts, and visual references), visual reference segmentation stands out for its unique flexibility and strong zero-shot capabilities. Recently, several SAM-based methods have made notable progress in this task by automatically generating prompts to guide SAM. However, these methods often generate prompts at boundaries of target regions due to suboptimal prompt encoder, which results in instability and reduced robustness. In this work, we introduce ProSAM, a simple but effective method to address the stability challenges we identified in existing SAM-based visual reference segmentation approaches. By learning a variational prompt encoder to predict multivariate prompt distributions, ProSAM avoids generating prompts that lie in unstable regions, overcoming the instability caused by less robust prompts. Our approach consistently surpasses state-of-the-art methods on the Pascal-5$^i$ and COCO-20$^i$ datasets, providing a more robust solution for visual reference segmentation.

视觉分割SAM改进概率提示鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。