SAM2到SAM3的分割范式断层:提示驱动失效,因模型转向概念理解。
The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
- 从空间提示转向多模态融合,实现语义感知与概念推理
- 新架构支持开放词汇、对比对齐和示例学习,突破几何限制
- 适合研究概念驱动图像分割的学者,或需泛化能力的视觉任务
本文剖析最新两代通用图像分割模型SAM2与SAM3之间的根本性断层。SAM2依赖点、框、掩码等空间提示进行纯几何与时间分割;而SAM3引入统一的视觉-语言架构,具备开放词汇推理、语义定位、对比对齐及基于示例的概念理解能力。分析从五个维度展开:(1) 提示型与概念型分割的本质差异;(2) 架构差异:SAM2为纯视觉-时序设计,SAM3集成视觉-语言编码器、几何与示例编码器、融合模块、DETR式解码器、对象查询及基于专家混合的歧义处理;(3) 数据集与标注差异:从SA-V视频掩码转为多模态概念标注语料库;(4) 训练与超参数差异,说明SAM2优化经验不适用于SAM3;(5) 评估体系变迁:由几何交并比(IoU)转向语义与开放词汇评价。这些分析确立SAM3为新一代分割基础模型,并指明概念驱动分割时代的发展方向。
原文摘要 · Abstract (English)
This paper investigates the fundamental discontinuity between the latest two Segment Anything Models: SAM2 and SAM3. We explain why the expertise in prompt-based segmentation of SAM2 does not transfer to the multimodal concept-driven paradigm of SAM3. SAM2 operates through spatial prompts points, boxes, and masks yielding purely geometric and temporal segmentation. In contrast, SAM3 introduces a unified vision-language architecture capable of open-vocabulary reasoning, semantic grounding, contrastive alignment, and exemplar-based concept understanding. We structure this analysis through five core components: (1) a Conceptual Break Between Prompt-Based and Concept-Based Segmentation, contrasting spatial prompt semantics of SAM2 with multimodal fusion and text-conditioned mask generation of SAM3; (2) Architectural Divergence, detailing pure vision-temporal design of SAM2 versus integration of vision-language encoders, geometry and exemplar encoders, fusion modules, DETR-style decoders, object queries, and ambiguity-handling via Mixture-of-Experts in SAM3; (3) Dataset and Annotation Differences, contrasting SA-V video masks with multimodal concept-annotated corpora of SAM3; (4) Training and Hyperparameter Distinctions, showing why SAM2 optimization knowledge does not apply to SAM3; and (5) Evaluation, Metrics, and Failure Modes, outlining the transition from geometric IoU metrics to semantic, open-vocabulary evaluation. Together, these analyses establish SAM3 as a new class of segmentation foundation model and chart future directions for the emerging concept-driven segmentation era.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。