基于全局与局部特征融合的胸部X光基础模型,提升多病种检测泛化能力。
Chest X-ray Foundation Model with Global and Local Representations Integration
- 采用全局与局部特征融合机制,增强疾病特异性识别能力。
- 在CXR-LT 24数据集上对40种病灶分类表现优于现有模型。
- 标签效率高,适用于小样本及分布外任务,适合临床部署。
胸部X光(CXR)是最常见的影像检查,广泛用于胸腔疾病检测和术后监测。然而,特定任务分类模型受限于范围、需大量标注数据且对分布外数据泛化能力差。为此,我们提出CheXFound,一个自监督视觉基础模型,通过在包含超过一百万张独特胸部X光片的CXR-1M数据集上预训练,学习鲁棒的图像表征,并在下游任务中实现良好泛化。我们设计了全局与局部表征融合(GLoRI)模块,结合疾病特异的局部特征与全局图像特征,提升多标签分类性能。实验表明,CheXFound在CXR-LT 24数据集上对40种不同流行程度的疾病发现分类表现优于现有模型,并在有限训练数据下展现出更强的标签效率。此外,其在分布外任务如机会性心血管风险评估和死亡率预测中也取得显著提升。这些结果凸显其强大泛化能力,支持多样化的高效适配。项目代码已公开:https://github.com/RPIDIAL/CheXFound。
原文摘要 · Abstract (English)
Chest X-ray (CXR) is the most frequently ordered imaging test, supporting diverse clinical tasks from thoracic disease detection to postoperative monitoring. However, task-specific classification models are limited in scope, require costly labeled data, and lack generalizability to out-of-distribution datasets. To address these challenges, we introduce CheXFound, a self-supervised vision foundation model that learns robust CXR representations and generalizes effectively across a wide range of downstream tasks. We pretrain CheXFound on a curated CXR-1M dataset, comprising over one million unique CXRs from publicly available sources. We propose a Global and Local Representations Integration (GLoRI) module for downstream adaptations, by incorporating disease-specific local features with global image features for enhanced performance in multilabel classification. Our experimental results show that CheXFound outperforms state-of-the-art models in classifying 40 disease findings across different prevalence levels on the CXR-LT 24 dataset and exhibits superior label efficiency on downstream tasks with limited training data. Additionally, CheXFound achieved significant improvements on new tasks with out-of-distribution datasets, including opportunistic cardiovascular disease risk estimation and mortality prediction. These results highlight CheXFound's strong generalization capabilities, enabling diverse adaptations with improved label efficiency. The project source code is publicly available at https://github.com/RPIDIAL/CheXFound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。