在胎儿脑超声图像中,通用模型难以区分结构相似的解剖平面,需专用预训练才能可靠识别。
Challenging DINOv3 Foundation Model under Low Inter-Class Variability: A Case Study on Fetal Brain Ultrasound
- 用18.8万张多中心胎儿脑超声图构建统一基准,评估DINOv3在低类间差异下的表现。
- 专用数据预训练使模型F1得分提升20%,显著优于自然图像初始化。
- 适合关注医学影像通用模型局限性与领域适配的研究者和临床开发者。
目的:本研究首次系统评估了基础模型在胎儿超声(US)成像中低类间差异条件下的表现。尽管近期视觉基础模型如DINOv3在医学领域展现出强大泛化能力,但其对解剖结构相似区域的判别能力尚未被系统研究。为此,我们聚焦于胎儿脑标准切面——丘脑水平(TT)、脑室水平(TV)和小脑水平(TC),这些切面具有高度重叠的解剖特征,对可靠的生物测量评估构成严峻挑战。方法:为确保公平可复现的评估,所有公开的胎儿超声数据集被整合为统一的多中心基准FetalUS-188K,包含超过18.8万张来自异构采集环境的标注图像。DINOv3采用自监督方式在该数据上预训练,以学习超声感知表征。随后通过标准化微调协议(冻结主干的线性探测与全量微调)在两种初始化方案下评估:(i) 在FetalUS-188K上预训练;(ii) 使用自然图像预训练的DINOv3权重初始化。结果:在胎儿超声数据上预训练的模型始终优于自然图像初始化模型,加权F1分数最高提升20%。领域自适应预训练使网络得以保留对区分中间切面(如TV)至关重要的细微回声与结构线索。结论:结果表明,通用基础模型在低类间差异下无法有效泛化,而领域特定预训练对于实现胎儿脑超声成像中稳健且临床可用的表示至关重要。
原文摘要 · Abstract (English)
Purpose: This study provides the first comprehensive evaluation of foundation models in fetal ultrasound (US) imaging under low inter-class variability conditions. While recent vision foundation models such as DINOv3 have shown remarkable transferability across medical domains, their ability to discriminate anatomically similar structures has not been systematically investigated. We address this gap by focusing on fetal brain standard planes--transthalamic (TT), transventricular (TV), and transcerebellar (TC)--which exhibit highly overlapping anatomical features and pose a critical challenge for reliable biometric assessment. Methods: To ensure a fair and reproducible evaluation, all publicly available fetal ultrasound datasets were curated and aggregated into a unified multicenter benchmark, FetalUS-188K, comprising more than 188,000 annotated images from heterogeneous acquisition settings. DINOv3 was pretrained in a self-supervised manner to learn ultrasound-aware representations. The learned features were then evaluated through standardized adaptation protocols, including linear probing with frozen backbone and full fine-tuning, under two initialization schemes: (i) pretraining on FetalUS-188K and (ii) initialization from natural-image DINOv3 weights. Results: Models pretrained on fetal ultrasound data consistently outperformed those initialized on natural images, with weighted F1-score improvements of up to 20 percent. Domain-adaptive pretraining enabled the network to preserve subtle echogenic and structural cues crucial for distinguishing intermediate planes such as TV. Conclusion: Results demonstrate that generic foundation models fail to generalize under low inter-class variability, whereas domain-specific pretraining is essential to achieve robust and clinically reliable representations in fetal brain ultrasound imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。