对比8种医学与通用大模型,发现医学预训练提升特征质量,但复杂病灶定位仍需微调。
Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- 比较8个医学与通用大模型在胸部X光上的表现,评估分类与分割任务
- 医学模型在分类和显著解剖结构分割中表现更优,但细微病灶定位差
- 无需文本对齐,监督学习模型在分割任务上仍可超越大模型
基础模型(FMs)有望泛化医学影像分析,但其效果因预训练领域(医学/通用)、范式(如文本引导)和架构而异。为探究此问题,我们评估了8个医学与通用领域基础模型的视觉编码器在胸部X光分析中的表现。通过线性探测和微调,在肺气肿、心脏扩大分类以及肺气肿、心边界分割任务上进行基准测试。结果表明,医学领域预训练具有显著优势:医学模型在线性探测中持续优于通用模型,具备更优的初始特征质量。然而特征效用高度依赖任务:预训练嵌入对全局分类和显著解剖结构分割(如心脏)有效;而对于复杂、细微病灶(如肺气肿)的分割,所有模型在未显著微调时表现均差,暴露出局部定位能力的关键短板。子组分析显示,模型在分类中使用混杂捷径(如胸管判断肺气肿),该策略在精确分割中失效。此外,昂贵的图文对齐并非必要;仅图像(RAD-DINO)和标签监督(Ark+)模型即位列前列。值得注意的是,监督端到端基线模型表现依然强劲,其分割性能匹配甚至超过最佳大模型。研究揭示:尽管医学预训练有益,但架构选择(如多尺度)至关重要,且预训练特征并非普遍有效,尤其在复杂定位任务中,监督模型仍是有力替代方案。
原文摘要 · Abstract (English)
Foundation models (FMs) promise to generalize medical imaging, but their effectiveness varies. It remains unclear how pre-training domain (medical vs. general), paradigm (e.g., text-guided), and architecture influence embedding quality, hindering the selection of optimal encoders for specific radiology tasks. To address this, we evaluate vision encoders from eight medical and general-domain FMs for chest X-ray analysis. We benchmark classification (pneumothorax, cardiomegaly) and segmentation (pneumothorax, cardiac boundary) using linear probing and fine-tuning. Our results show that domain-specific pre-training provides a significant advantage; medical FMs consistently outperformed general-domain models in linear probing, establishing superior initial feature quality. However, feature utility is highly task-dependent. Pre-trained embeddings were strong for global classification and segmenting salient anatomy (e.g., heart). In contrast, for segmenting complex, subtle pathologies (e.g., pneumothorax), all FMs performed poorly without significant fine-tuning, revealing a critical gap in localizing subtle disease. Subgroup analysis showed FMs use confounding shortcuts (e.g., chest tubes for pneumothorax) for classification, a strategy that fails for precise segmentation. We also found that expensive text-image alignment is not a prerequisite; image-only (RAD-DINO) and label-supervised (Ark+) FMs were among top performers. Notably, a supervised, end-to-end baseline remained highly competitive, matching or exceeding the best FMs on segmentation tasks. These findings show that while medical pre-training is beneficial, architectural choices (e.g., multi-scale) are critical, and pre-trained features are not universally effective, especially for complex localization tasks where supervised models remain a strong alternative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。