arXiv:2512.24492eess.IVcs.AI2025-12被引 1

用自监督学习提升早孕胎儿心脏超声图像分类准确率

Automated Classification of First-Trimester Fetal Heart Views Using Ultrasound-Specific Self-Supervised Learning

  • 基于37万张无标注超声图预训练,用掩码自编码方法构建基础模型
  • 在6720张早孕超声图上分类5类视图,准确率达90.57%,优于现有模型
  • 无需复杂图像处理,能更好识别无效帧,适合临床辅助诊断

先天性心脏病是最常见的先天畸形,也是新生儿发病和死亡的主要原因。尽管早孕期胎儿超声心动图提供了早期检测机会,但因心脏结构小、信噪比低及操作者差异大,自动化分析仍具挑战。本文评估了一种针对超声的自监督基础模型USF-MAE,用于早孕胎儿心脏视图分类。USF-MAE在超过37万张涵盖40多个解剖区域的未标注超声图像上通过掩码自编码进行预训练,随后在公开数据集上微调,该数据集包含6,720张早孕胎儿超声心动图,需分类为五类:主动脉、房室血流、V征、X征及其他。模型性能与监督式卷积神经网络(ResNet-18、ResNet-50)及在自然图像上预训练的ViT-B/16对比,所有模型采用相同预处理、数据划分与优化协议。在独立测试集上,USF-MAE在各项指标中表现最优,准确率90.57%、精确率91.15%、召回率90.57%、F1分数90.71%,相比最强基线ResNet-18分别提升+2.03%准确率和+1.98% F1分数。该方法无需激进图像预处理或感兴趣区域裁剪,且对非诊断性帧的判别能力更强。

原文摘要 · Abstract (English)

Congenital heart disease remains the most common congenital anomaly and a leading cause of neonatal morbidity and mortality. Although first-trimester fetal echocardiography offers an opportunity for earlier detection, automated analysis at this stage is challenging due to small cardiac structures, low signal-to-noise ratio, and substantial inter-operator variability. In this work, we evaluate a self-supervised ultrasound foundation model, USF-MAE, for first-trimester fetal heart view classification. USF-MAE is pretrained using masked autoencoding modelling on more than 370,000 unlabelled ultrasound images spanning over 40 anatomical regions and is subsequently fine-tuned for downstream classification. As a proof of concept, the pretrained Vision Transformer encoder was fine-tuned on an open-source dataset of 6,720 first-trimester fetal echocardiography images to classify five categories: aorta, atrioventricular flows, V sign, X sign, and Other. Model performance was benchmarked against supervised convolutional neural network baselines (ResNet-18 and ResNet-50) and a Vision Transformer (ViT-B/16) model pretrained on natural images (ImageNet-1k). All models were trained and evaluated using identical preprocessing, data splits, and optimization protocols. On an independent test set, USF-MAE achieved the highest performance across all evaluation metrics, with 90.57% accuracy, 91.15% precision, 90.57% recall, and 90.71% F1-score. This represents an improvement of +2.03% in accuracy and +1.98% in F1-score compared with the strongest baseline, ResNet-18. The proposed approach demonstrated robust performance without reliance on aggressive image preprocessing or region-of-interest cropping and showed improved discrimination of non-diagnostic frames.

医学影像自监督学习胎儿超声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。