对比多种模型在跨国胸片数据上的诊断能力,发现视觉语言模型表现更优。
Benchmarking Chest X-ray Diagnosis Models Across Multinational Datasets
- 用多国数据集对比8种模型,涵盖视觉语言模型与传统CNN
- 知识增强提示+结构化监督的MAVL模型在公开/私有数据上均领先
- 儿童胸片诊断性能显著下降,提示需针对性优化
基于视觉-语言预训练的基础模型在胸片(CXR)解读中展现出潜力,但其在多样人群和诊断任务中的实际表现仍缺乏充分评估。本研究在来自美国、西班牙、印度和越南的6个公开数据集及中国三家医院的3个私有数据集上,对8种胸片诊断模型(5个视觉语言基础模型和3个基于CNN的架构)进行基准测试。评估涵盖37项标准化分类任务,使用AUROC、AUPRC等指标衡量性能。结果表明,基础模型在准确率和任务覆盖范围上均优于传统CNN。其中,采用知识增强提示和结构化监督的MAVL模型在公开数据集(平均AUROC: 0.82;AUPRC: 0.32)和私有数据集(平均AUROC: 0.95;AUPRC: 0.89)上表现最佳,在37项公开任务中排名第一14次,4项私有任务中排名第一3次。所有模型在儿童病例上性能均下降,平均AUROC从成人组的0.88 ± 0.18降至儿童组的0.57 ± 0.29(p = 0.0202)。研究强调了结构化监督与提示设计在放射学AI中的价值,并建议未来方向包括地理扩展与集成建模以支持临床部署。所有模型代码已开源。
原文摘要 · Abstract (English)
Foundation models leveraging vision-language pretraining have shown promise in chest X-ray (CXR) interpretation, yet their real-world performance across diverse populations and diagnostic tasks remains insufficiently evaluated. This study benchmarks the diagnostic performance and generalizability of foundation models versus traditional convolutional neural networks (CNNs) on multinational CXR datasets. We evaluated eight CXR diagnostic models - five vision-language foundation models and three CNN-based architectures - across 37 standardized classification tasks using six public datasets from the USA, Spain, India, and Vietnam, and three private datasets from hospitals in China. Performance was assessed using AUROC, AUPRC, and other metrics across both shared and dataset-specific tasks. Foundation models outperformed CNNs in both accuracy and task coverage. MAVL, a model incorporating knowledge-enhanced prompts and structured supervision, achieved the highest performance on public (mean AUROC: 0.82; AUPRC: 0.32) and private (mean AUROC: 0.95; AUPRC: 0.89) datasets, ranking first in 14 of 37 public and 3 of 4 private tasks. All models showed reduced performance on pediatric cases, with average AUROC dropping from 0.88 +/- 0.18 in adults to 0.57 +/- 0.29 in children (p = 0.0202). These findings highlight the value of structured supervision and prompt design in radiologic AI and suggest future directions including geographic expansion and ensemble modeling for clinical deployment. Code for all evaluated models is available at https://drive.google.com/drive/folders/1B99yMQm7bB4h1sVMIBja0RfUu8gLktCE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。