评估SAM3在遥感图像中的零样本与单样本分割能力,发现视觉提示更优。
Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing

- 将SAM3的二元存在头改造为零样本分类器,实现多任务评估
- 视觉提示可精准对齐遥感几何结构,文本提示引入语义偏差导致坐标回归下降
- 无需训练的代理评估协议,揭示模型在遥感场景下的局限与优化方向
大规模基础模型如分割一切模型3(SAM 3)有望推动开放词汇、无训练计算机视觉的发展。然而,其在地球观测影像复杂俯视几何结构上的分布外泛化能力尚未被量化。针对SAM 3在专业领域表现差异,我们开展了一项全面的多任务实证评估,涵盖遥感场景分类、目标检测和实例分割,在严格的零样本与单样本约束下进行。为此,我们通过重构SAM 3的解耦二元存在头,将其转为独立的零样本分类器;并通过五种配置系统分离文本与视觉提示模态,明确诊断模型多模态解码器中的对齐机制。结果表明:视觉提示能有效对齐复杂遥感几何结构,而文本提示则引入地表级语义偏见,显著干扰坐标回归。为避免资源密集型训练,我们提出一种新型无训练代理评估协议,用于广义零样本任务(场景分类与实例分割)。最终结果显示,SAM 3避免了传统领域自适应模型的过拟合问题,在分割任务中取得高调和均值分数。但其仍受亚像素分辨率限制与语义盲区影响,明确要求对多模态解码器进行参数高效地理空间微调。
原文摘要 · Abstract (English)
The deployment of large-scale foundation models, such as the Segment Anything Model 3 (SAM 3), promises a transition toward open-vocabulary, training-free computer vision. However, their capacity to generalize out-of-distribution to the complex, top-down geometric structures of Earth Observation imagery remains largely unquantified. Driven by SAM 3's performance disparities in highly specialized domains, we present a comprehensive, multi-task empirical evaluation across remote sensing scene classification, object detection, and instance segmentation under strict zero-shot and one-shot constraints. To achieve this, we introduce a structural adaptation of SAM 3 by repurposing its decoupled binary presence head into a standalone zero-shot classifier. Furthermore, by systematically isolating textual and visual prompt modalities across five configurations, we explicitly diagnose the alignment mechanics within the model's multimodal decoder. Our findings reveal severe cross-modal interference: while visual prompts successfully align the decoder to complex remote sensing geometry, textual prompts inject misaligned, ground-level semantic bias, actively degrading coordinate regression. To benchmark these capabilities without resource-intensive training, we formulate a novel training-free proxy evaluation protocol for Generalized Zero-Shot tasks (scene classification and instance segmentation). Ultimately, our results demonstrate that SAM 3 avoids the overfitting commonly seen in legacy domain-adapted models, achieving high Harmonic Mean scores in segmentation tasks. However, it remains fundamentally constrained by sub-pixel resolution limits and overhead semantic blind spots, charting a definitive mandate for parameter-efficient geospatial fine-tuning of its multimodal decoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。