首个面向胎儿超声的区域可控多模态模型,提升局部结构识别能力。
SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis

- 用分割掩码作为视觉提示,实现全局与局部特征联合学习。
- 在144万张图像上预训练,零样本迁移性能超越现有方法。
- 适合临床医生和研究者用于胎儿超声自动化分析。
视觉-语言基础模型在医学图像分析中展现出巨大潜力。尽管超声成像的基础模型已出现,但该领域仍具挑战性,主要由于严重散斑噪声、成像变异性和细微解剖边界,导致观察者间差异大。现有基于CLIP的模型主要依赖全局图像-文本对齐,难以捕捉临床关键局部结构。我们提出SonoCLIP,首个百万级可控制区域的胎儿超声视觉-语言基础模型,将分割掩码作为掩码通道视觉提示嵌入视觉编码器,实现全局-局部联合对比表征学习。为支持大规模区域-文本对齐,引入基于sigmoid的成对对比损失,增强大规模监督下的稳定性。我们进一步构建了包含144万张图像的多模态胎儿超声数据集,覆盖24个标准切面,用于大规模预训练。跨中心实验证明,SonoCLIP在全局与掩码引导推理下均实现优越的零样本迁移性能,建立了一个可控且面向临床的胎儿超声分析基础模型。代码与数据可在https://github.com/Harrison-one/SonoCLIP获取。
原文摘要 · Abstract (English)
Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomical boundaries, leading to high inter-observer variability. Existing CLIP-based models rely primarily on global image-text alignment, limiting their sensitivity to clinically decisive local structures. We propose SonoCLIP, the first million-scale region-controllable fetal ultrasound vision-language foundation model that integrates segmentation masks as mask-channel visual prompts within the vision encoder, enabling joint global-local contrastive representation learning. To support scalable region-text alignment, we introduce a sigmoid-based pairwise contrastive loss that improves stability under large-scale supervision. We further curate a 1.44M-image multimodal fetal ultrasound dataset spanning 24 standard planes for large-scale pretraining. Extensive cross-center evaluations demonstrate that SonoCLIP achieves superior zero-shot transfer performance under both global and mask-guided inference, establishing a controllable and clinically oriented foundation model for fetal ultrasound analysis. Our code and data are available at https://github.com/Harrison-one/SonoCLIP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。