首个专用于胎儿超声的多模态基础模型,实现精准图像理解与少样本泛化。
FetalCLIP: A Visual-Language Foundation Model for Fetal Ultrasound Image Analysis
- 基于21万张配对超声图与文本预训练,学习胎儿解剖特征。
- 在多种任务中超越基线,少样本下仍保持高精度。
- 适合医学影像研究者、产科智能诊断开发者使用。
基础模型在医疗领域日益有效,能通过大规模数据预训练快速适配下游任务。然而,胎儿超声因图像复杂且配对多模态数据稀缺,仍是基础模型应用的难点。为此,我们提出FetalCLIP,一种视觉-语言基础模型,可生成胎儿超声图像的通用表征。该模型在包含210,035张胎儿超声图像与对应文本的多样化数据集上进行多模态预训练,是迄今同类模型中规模最大的配对数据集。这一独特训练方式使FetalCLIP能有效学习胎儿超声中的复杂解剖特征,生成鲁棒表征,适用于多种下游任务。在分类、孕周估计、先天性心脏病(CHD)检测和胎儿结构分割等关键任务上的广泛基准测试表明,FetalCLIP全面优于现有基线,展现出卓越泛化能力,即使在少量标注数据下也表现优异。我们将公开发布FetalCLIP模型,以促进科学界发展。
原文摘要 · Abstract (English)
Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging domain for foundation models due to their inherent complexity, often requiring substantial additional training and facing limitations due to the scarcity of paired multimodal data. To overcome these challenges, here we introduce FetalCLIP, a vision-language foundation model capable of generating universal representation of fetal ultrasound images. FetalCLIP was pre-trained using a multimodal learning approach on a diverse dataset of 210,035 fetal ultrasound images paired with text. This represents the largest paired dataset of its kind used for foundation model development to date. This unique training approach allows FetalCLIP to effectively learn the intricate anatomical features present in fetal ultrasound images, resulting in robust representations that can be used for a variety of downstream applications. In extensive benchmarking across a range of key fetal ultrasound applications, including classification, gestational age estimation, congenital heart defect (CHD) detection, and fetal structure segmentation, FetalCLIP outperformed all baselines while demonstrating remarkable generalizability and strong performance even with limited labeled data. We plan to release the FetalCLIP model publicly for the benefit of the broader scientific community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。