对比两款胸部X光大模型表现,发现性能与跨数据集稳定性各有优劣。
Benchmarking CXR Foundation Models With Publicly Available MIMIC-CXR and NIH-CXR14 Datasets
- 统一预处理和分类器,公平比较两个大模型在两种公开数据集上的表现。
- MedImageInsight整体性能略高,而CXR-Foundation跨数据集更稳定。
- 结果可复现,为后续医疗多模态研究提供基准参考。
近期基础模型在医学图像表征学习中表现优异,但其在不同数据集上的对比行为仍不明确。本文在公开的MIMIC-CXR和NIH ChestX-ray14数据集上,对两个大规模胸部X光(CXR)嵌入模型(CXR-Foundation(ELIXR v2.0)和MedImageInsight)进行基准测试。每个模型均采用统一预处理流程和固定下游分类器,确保可复现性。直接从预训练编码器提取嵌入,使用轻量级LightGBM分类器在多个疾病标签上训练,并报告均值AUROC与F1-score及其95%置信区间。结果表明,MedImageInsight在多数任务中表现略优,而CXR-Foundation展现出更强的跨数据集稳定性。对MedImageInsight嵌入的无监督聚类进一步揭示了与定量结果一致的疾病特异性结构。研究强调了标准化评估的重要性,并为未来多模态及临床整合研究建立了可复现基线。
原文摘要 · Abstract (English)
Recent foundation models have demonstrated strong performance in medical image representation learning, yet their comparative behaviour across datasets remains underexplored. This work benchmarks two large-scale chest X-ray (CXR) embedding models (CXR-Foundation (ELIXR v2.0) and MedImagelnsight) on public MIMIC-CR and NIH ChestX-ray14 datasets. Each model was evaluated using a unified preprocessing pipeline and fixed downstream classifiers to ensure reproducible comparison. We extracted embeddings directly from pre-trained encoders, trained lightweight LightGBM classifiers on multiple disease labels, and reported mean AUROC, and F1-score with 95% confidence intervals. MedImageInsight achieved slightly higher performance across most tasks, while CXR-Foundation exhibited strong cross-dataset stability. Unsupervised clustering of MedImageIn-sight embeddings further revealed a coherent disease-specific structure consistent with quantitative results. The results highlight the need for standardised evaluation of medical foundation models and establish reproducible baselines for future multimodal and clinical integration studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。