arXiv:2507.09562cs.CVcs.AI2025-07被引 5

系统梳理SAM模型的提示工程方法与应用挑战

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

  • 按几何、语义、多模态分类提示工程方法
  • 实现医学影像、遥感等跨领域零样本分割
  • 适合研究大模型提示设计与视觉任务的学者

Segment Anything Model(SAM)通过提示驱动范式实现了强大的零样本图像分割能力。提示作为人意图与机器感知之间的语义接口,其工程设计直接影响模型性能。本文首次系统综述了SAM及其生态中提示工程的发展现状。提出层次化分类体系,将方法分为几何提示、文本语义提示和多模态融合提示,分析其设计原则与应用目标。进一步考察从人工设计到基于检测器输出、原型学习、强化学习及视觉语言模型的自动化演进。揭示提示工程如何推动SAM在医学影像、遥感、工业质检与异常检测等领域的泛化能力。识别出提示敏感性、跨模态错位与计算效率等关键挑战,并展望因果提示推理、多智能体协作提示与扩散渐进优化等前沿方向。本综述为理解分割基础模型中的提示工程提供统一视角,助力未来研究。

原文摘要 · Abstract (English)

The Segment Anything Model (SAM) has transformed image segmentation by introducing a prompt-based paradigm that enables strong zero-shot generalization. In this framework, prompts serve as a semantic interface between human intent and machine perception, making prompt engineering a central factor in model performance. Despite its importance, prompt engineering within SAM and its variants has not yet been systematically reviewed in the literature. This survey addresses that gap by providing a structured and comprehensive overview of prompt engineering techniques developed for SAM and its rapidly growing ecosystem. We introduce a hierarchical taxonomy that organizes methods into geometric prompts, textual semantic prompts, and multimodal fusion prompts, and analyze how these categories reflect different design principles and application goals. In addition, we examine the transition from manually crafted prompts to more advanced, automated approaches based on detector outputs, prototype learning, reinforcement learning, and vision-language models. Beyond categorizing existing work, we trace how prompt engineering has enabled SAM to generalize across domains such as medical imaging, remote sensing, industrial inspection, and anomaly detection. We further identify key challenges---including prompt sensitivity, cross-modal misalignment, and computational inefficiency---and highlight promising research directions such as causal prompt reasoning, collaborative multi-agent prompting, and diffusion-based progressive refinement. By consolidating these developments into a unified perspective, our survey provides a timely reference for understanding the role of prompt engineering in segmentation foundation models and lays the groundwork for future advances in this evolving field.

提示工程图像分割SAM大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。