arXiv:2510.22990eess.IVcs.AI2025-10被引 15

首个专为超声设计的自监督大模型,用无标签数据学出强泛化能力。

USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoencoding

  • 用掩码自编码重建超声图像块,从37万张无标签图中自学特征。
  • 在3个肿瘤分类任务中准确率超80%,接近有监督模型表现。
  • 适合医疗影像研究者,尤其缺标注数据的超声领域。

超声成像是广泛应用的诊断手段,具有实时、无辐射等优势,但受噪声高、操作者依赖和视野有限影响,判读难度大,存在显著观察者间差异。当前深度学习方法受限于大规模标注数据稀缺及通用图像与超声图像间的领域差异,导致预训练模型迁移效果差。为此,我们提出首个仅在超声数据上预训练的大规模自监督基础模型USF-MAE,采用掩码自编码框架。模型在涵盖46个开源数据集的37万张2D/3D超声图像(统称OpenUS-46)上预训练,覆盖二十多个解剖区域,该数据集已公开。基于视觉变换器架构,模型通过重建被遮蔽图像块,学习到丰富的模态特异性表示。在三个公开下游分类任务(BUS-BRA乳腺癌、MMOTU-2D卵巢肿瘤、GIST514-DB胃肠道间质瘤)上微调,分别取得81.6%、79.6%、82.4%的F1分数,均优于传统CNN和ViT基线。尽管预训练阶段未使用标签,其性能接近监督型基础模型UltraSam,在乳腺癌任务上相当,其余任务更优,展现强大跨解剖区域泛化能力。

原文摘要 · Abstract (English)

Ultrasound imaging is one of the most widely used diagnostic modalities, offering real-time, radiation-free assessment across diverse clinical domains. However, interpretation of ultrasound images remains challenging due to high noise levels, operator dependence, and limited field of view, resulting in substantial inter-observer variability. Current Deep Learning approaches are hindered by the scarcity of large labeled datasets and the domain gap between general and sonographic images, which limits the transferability of models pretrained on non-medical data. To address these challenges, we introduce the Ultrasound Self-Supervised Foundation Model with Masked Autoencoding (USF-MAE), the first large-scale self-supervised MAE framework pretrained exclusively on ultrasound data. The model was pre-trained on 370,000 2D and 3D ultrasound images curated from 46 open-source datasets, collectively termed OpenUS-46, spanning over twenty anatomical regions. This curated dataset has been made publicly available to facilitate further research and reproducibility. Using a Vision Transformer encoder-decoder architecture, USF-MAE reconstructs masked image patches, enabling it to learn rich, modality-specific representations directly from unlabeled data. The pretrained encoder was fine-tuned on three public downstream classification benchmarks: BUS-BRA (breast cancer), MMOTU-2D (ovarian tumors), and GIST514-DB (gastrointestinal stromal tumors). Across all tasks, USF-MAE consistently outperformed conventional CNN and ViT baselines, achieving F1-scores of 81.6%, 79.6%, and 82.4%, respectively. Despite not using labels during pretraining, USF-MAE approached the performance of the supervised foundation model UltraSam on breast cancer classification and surpassed it on the other tasks, demonstrating strong cross-anatomical generalization.

超声影像自监督学习基础模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。