arXiv:2601.10880cs.CVcs.AI2026-01被引 21

医学影像分割新基座,支持文本提示的通用模型

Medical SAM3: A Foundation Model for Universal Prompt-Driven Medical Image Segmentation

  • 全量微调SAM3,在33个医疗数据集上学习跨模态语义
  • 在复杂解剖结构中表现显著提升,3D长程上下文任务准确率提高21%
  • 适合需要精准、灵活分割的临床研究与医学AI开发

可提示分割基础模型如SAM3通过交互式和概念式提示展现出强大泛化能力。然而,其在医学图像分割中的直接应用受限于严重的领域偏移、缺乏特权空间提示,以及对复杂解剖与体积结构的推理需求。本文提出Medical SAM3,一个面向通用提示驱动医学图像分割的基础模型,通过在大规模异构二维与三维医学影像数据集上对SAM3进行全量微调,并配以分割掩码与文本提示。系统分析表明,原始SAM3在医学数据上性能显著下降,其竞争力主要依赖于地面真值生成的边界框等强几何先验。这一发现促使我们超越仅靠提示工程的改进,实施全模型适应。在涵盖10种医学成像模态的33个数据集上微调后,Medical SAM3获得稳健的领域特异性表征,同时保持提示驱动的灵活性。跨器官、模态与维度的大量实验显示,其在语义模糊、形态复杂及长程3D上下文等挑战性场景中持续实现显著性能提升。结果确立了Medical SAM3作为医学影像通用文本引导分割基础模型的地位,并强调了在严重领域偏移下实现鲁棒提示驱动分割需进行整体模型适配。代码与模型将公开于 https://github.com/AIM-Research-Lab/Medical-SAM3。

原文摘要 · Abstract (English)

Promptable segmentation foundation models such as SAM3 have demonstrated strong generalization capabilities through interactive and concept-based prompting. However, their direct applicability to medical image segmentation remains limited by severe domain shifts, the absence of privileged spatial prompts, and the need to reason over complex anatomical and volumetric structures. Here we present Medical SAM3, a foundation model for universal prompt-driven medical image segmentation, obtained by fully fine-tuning SAM3 on large-scale, heterogeneous 2D and 3D medical imaging datasets with paired segmentation masks and text prompts. Through a systematic analysis of vanilla SAM3, we observe that its performance degrades substantially on medical data, with its apparent competitiveness largely relying on strong geometric priors such as ground-truth-derived bounding boxes. These findings motivate full model adaptation beyond prompt engineering alone. By fine-tuning SAM3's model parameters on 33 datasets spanning 10 medical imaging modalities, Medical SAM3 acquires robust domain-specific representations while preserving prompt-driven flexibility. Extensive experiments across organs, imaging modalities, and dimensionalities demonstrate consistent and significant performance gains, particularly in challenging scenarios characterized by semantic ambiguity, complex morphology, and long-range 3D context. Our results establish Medical SAM3 as a universal, text-guided segmentation foundation model for medical imaging and highlight the importance of holistic model adaptation for achieving robust prompt-driven segmentation under severe domain shift. Code and model will be made available at https://github.com/AIM-Research-Lab/Medical-SAM3.

医学影像分割模型提示驱动SAM3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。