arXiv:2607.28771cs.CV2026-07

评估医学大模型在非洲脑影像上的泛化能力,发现数据量是关键瓶颈。

Do Medical Foundation Models Generalize on the African Brain?

论文配图:Do Medical Foundation Models Generalize on the African Brain?
图 1 · 摘自论文原文
  • 对比多种大模型在尼日利亚和非非洲数据上表现,测试其跨区域泛化能力。
  • 分类任务中最高准确率仅0.86 AUC,分割任务提升至0.86 Dice,但差异不显著。
  • 研究揭示非洲数据稀缺才是主要障碍,而非模型本身存在偏见。

医学基础模型(FMs)在脑部MRI分析中应用日益广泛,但其评估仍以高资源数据集为主,对非洲人群的泛化能力研究不足。本文通过两个任务评估:使用尼日利亚数据集进行痴呆分类,以及使用BraTS-Africa数据集进行脑肿瘤分割。比较了两种通用型模型(BrainIAC、3DINO)和两种专用分割模型(MedSAM2、Medical-SAM2)与从零训练基线的性能。在分类任务中,模型提升有限(最高ROC-AUC为0.86,由BrainIAC实现);而在分割任务中,模型持续提升性能,最高达0.86 Dice(MedSAM2)。非洲与非非洲数据集间的表现差异不一致,且更可能受数据规模影响而非数据来源。结果表明,基础模型并无对非洲人群的固有偏见,而非洲神经影像数据的匮乏与多样性不足是阻碍模型可靠评估与部署的主要原因。

原文摘要 · Abstract (English)

Medical foundation models (FMs) are increasingly used for brain MRI analysis. However, their evaluation remains dominated by high-resource datasets, leaving generalization to African cohorts underexplored. We assess whether FMs generalize equally to African and non-African brain MRI data across two tasks: dementia classification using a Nigerian dataset and brain tumor segmentation using BraTS-Africa. We evaluate two generalist FMs (BrainIAC, 3DINO) and two segmentation-specific FMs (MedSAM2, Medical-SAM2) against a from-scratch baseline. For classification, FMs provide limited gains (highest ROC-AUC of 0.86 with BrainIAC), whereas for segmentation they consistently improve performance, reaching up to 0.86 Dice with MedSAM2. Performance differences between African and non-African cohorts are inconsistent and appear more related to dataset size than data origin. These results suggest that FMs do not exhibit an inherent bias against African cohorts, and highlight the limited availability and diversity of African neuroimaging datasets as the main barrier to robust evaluation and deployment.

医学AI脑影像泛化性非洲数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。