arXiv:2410.22048cs.CV2024-10被引 8

对比人类与自动提示在图像分割中的表现,发现人类更优且可提升自动方法。

Benchmarking Human and Automated Prompting in the Segment Anything Model

  • 基于真实数据集对比人类与自动点提示策略
  • 人类提示分割准确率比自动方法高29%
  • 通过微调可使自动方法性能提升68%,适合优化提示设计

Segment Anything Model(SAM)在直观交互式图像分割任务中表现出色,激发了对高效视觉提示设计的兴趣。尽管已有自动化点提示选择策略,但其效果与人类相比如何、在不同图像领域表现如何,仍缺乏深入理解。此外,自动化提示在SAM微调中的作用及提示点间距离等可解释因素对分割性能的影响也未被探索。为此,我们利用新发布的视觉提示数据集PointPrompt,设计了一系列基准任务,以增进对人类提示与自动化提示差异的理解,并识别有效提示的关键因素。实验表明,人类提示的分割得分比自动化策略高出约29%,并发现若干特征可预测提示性能($R^2 > 0.5$)。同时,通过微调可使自动化方法性能提升高达68%。研究揭示了人机提示间的差距,并指明了改进提示设计的可行路径。更多细节、数据集链接与代码见https://github.com/olivesgatech/PointPrompt。

原文摘要 · Abstract (English)

The remarkable capabilities of the Segment Anything Model (SAM) for tackling image segmentation tasks in an intuitive and interactive manner has sparked interest in the design of effective visual prompts. Such interest has led to the creation of automated point prompt selection strategies, typically motivated from a feature extraction perspective. However, there is still very little understanding of how appropriate these automated visual prompting strategies are, particularly when compared to humans, across diverse image domains. Additionally, the performance benefits of including such automated visual prompting strategies within the finetuning process of SAM also remains unexplored, as does the effect of interpretable factors like distance between the prompt points on segmentation performance. To bridge these gaps, we leverage a recently released visual prompting dataset, PointPrompt, and introduce a number of benchmarking tasks that provide an array of opportunities to improve the understanding of the way human prompts differ from automated ones and what underlying factors make for effective visual prompts. We demonstrate that the resulting segmentation scores obtained by humans are approximately 29% higher than those given by automated strategies and identify potential features that are indicative of prompting performance with $R^2$ scores over 0.5. Additionally, we demonstrate that performance when using automated methods can be improved by up to 68% via a finetuning approach. Overall, our experiments not only showcase the existing gap between human prompts and automated methods, but also highlight potential avenues through which this gap can be leveraged to improve effective visual prompt design. Further details along with the dataset links and codes are available at https://github.com/olivesgatech/PointPrompt

图像分割提示工程人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。