全面评估SAM系列对复杂上下文概念的分割能力,揭示其在真实场景中的表现边界。
Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes
- 构建统一评测框架,支持手动、自动与混合提示,覆盖多模态场景
- 在11类上下文依赖概念上测试,发现模型对医学病变等任务表现不足
- 提出提示鲁棒性测试,模拟真实世界不完美输入,助力未来模型优化
基于数十亿图像-掩码对训练的SAM及其升级版SAM~2,在计算机视觉多个领域产生深远影响。凭借前所未有的数据多样性,它们展现出强大的开集分割能力,其中SAM~2进一步支持高质量视频分割。尽管在人、车、道路等独立于上下文的概念上表现优异,但对视觉显著性、伪装、工业缺陷和医学病灶等依赖上下文的概念仍存在识别短板。这些概念高度依赖全局与局部上下文信息,易受场景变化影响,需模型具备强判别能力。现有对SAM系列的评估尚不全面,限制了对其性能边界的理解,也阻碍未来模型设计。本文在自然、医疗与工业场景中,针对2D图像、3D图像及视频,对11类上下文依赖概念进行系统评测。构建统一评估框架,支持人工、自动及中间自提示,并结合特定提示生成与交互策略。探索SAM~2在上下文学习中的潜力,引入提示鲁棒性测试以模拟现实中的不完美提示。分析SAM系列在理解上下文依赖概念上的优劣,讨论其在分割任务中的未来发展路径。
原文摘要 · Abstract (English)
As large-scale foundation models trained on billions of image--mask pairs covering a vast diversity of scenes, objects, and contexts, SAM and its upgraded version, SAM~2, have significantly influenced multiple fields within computer vision. Leveraging such unprecedented data diversity, they exhibit strong open-world segmentation capabilities, with SAM~2 further enhancing these capabilities to support high-quality video segmentation. While SAMs (SAM and SAM~2) have demonstrated excellent performance in segmenting context-independent concepts like people, cars, and roads, they overlook more challenging context-dependent (CD) concepts, such as visual saliency, camouflage, industrial defects, and medical lesions. CD concepts rely heavily on global and local contextual information, making them susceptible to shifts in different contexts, which requires strong discriminative capabilities from the model. The lack of comprehensive evaluation of SAMs limits understanding of their performance boundaries, which may hinder the design of future models. In this paper, we conduct a thorough evaluation of SAMs on 11 CD concepts across 2D and 3D images and videos in various visual modalities within natural, medical, and industrial scenes. We develop a unified evaluation framework for SAM and SAM~2 that supports manual, automatic, and intermediate self-prompting, aided by our specific prompt generation and interaction strategies. We further explore the potential of SAM~2 for in-context learning and introduce prompt robustness testing to simulate real-world imperfect prompts. Finally, we analyze the benefits and limitations of SAMs in understanding CD concepts and discuss their future development in segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。