让超声图像分割支持自然语言指令,通用性强且无需专家提示。
UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

- 基于超声特有图像-掩码-概念三元组训练,实现文本驱动分割。
- 在37个数据集、13类器官上训练,跨器官分割性能领先。
- 内置指令解析代理,可处理复杂自然语言查询,适合临床使用。
超声成像因便携、低成本和实时性广泛应用于临床,其图像分割至关重要。但超声图像常受斑点噪声、低对比度、声影和边界模糊影响,与CT、MRI等模态差异显著。现有方法多为特定任务模型或依赖专家提供视觉提示的基座模型,难以灵活应用。为此,我们提出UltraSAM3——一种面向通用超声图像分割的概念驱动基座模型。该模型通过将SAM3适配至超声特有的图像-掩码-概念三元组,实现文本驱动的目标指定。模型在覆盖37个公开数据集和13个解剖类别、大规模的超声分割语料库上训练,能够对不同器官和病灶的临床有意义概念进行视觉模式对齐。为进一步提升真实临床交互下的可用性,我们设计了一种指令引导代理,可将复杂自然语言查询解析为简洁的超声概念提示。大量实验表明,UltraSAM3在多器官超声基准测试、外部数据集及视觉提示增强设置下均优于代表性概念与文本驱动生物医学分割模型。此外,该代理显著提升了复杂用户指令下的分割鲁棒性。结果表明,超声特异性概念适配对于构建可泛化、可交互的超声分割基座模型是有效的。
原文摘要 · Abstract (English)
Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound image segmentation important. However, ultrasound images differ substantially from CT, MRI, and other medical imaging modalities, as they are often affected by speckle noise, low contrast, acoustic shadows and ambiguous boundaries. Existing ultrasound segmentation methods are still mainly limited to task-specific models or visual-prompt-based foundation models, which are either tailored to particular tasks or require expert-provided visual prompts, making them inconvenient for flexible clinical use. To address these challenges, we propose UltraSAM3, a concept-driven foundation model for universal ultrasound image segmentation. Unlike conventional models, UltraSAM3 enables text-based target specification by adapting SAM3 to ultrasound-specific image--mask--concept triplets. The model is trained on a large-scale ultrasound segmentation corpus covering 37 public datasets and 13 anatomical categories, allowing it to align ultrasound visual patterns with clinically meaningful concepts across diverse organs and lesions. To further improve usability under realistic clinical interaction, we propose an instruction-guided agent that parses complex natural language queries into concise ultrasound concept prompts for UltraSAM3. Extensive experiments demonstrate that UltraSAM3 consistently outperforms representative concept- and text-driven biomedical segmentation models on multi-organ ultrasound benchmarks, external datasets, and visual-prompt-enhanced settings. Moreover, the agent improves segmentation robustness for complex user instructions. These results indicate that ultrasound-specific concept adaptation is effective for building generalizable and interactive ultrasound segmentation foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。