公开胸片数据集存在标注质量差、分布偏移和评估不公问题,影响AI模型临床可信度。
Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
- 通过多模型跨数据集测试发现外部性能显著下降,AUPRC与F1分数大幅降低
- 专家评审显示标注与真实诊断差异显著,自动化标签存在误判风险
- 模型能精准区分不同数据集,暴露年龄性别群体的性能偏差问题
人工智能在胸部X光影像分析中已展现出接近放射科医生的诊断能力,这得益于MIMIC-CXR、ChestX-ray14、PadChest和CheXpert等大型公开数据集提供的数十万张带病理标注的图像。然而这些数据集存在重要局限:从报告自动提取标签引入错误,尤其在处理不确定性和否定句时;放射科医生审核常与标注结果不一致。此外,领域偏移和人群偏差限制了模型泛化能力,而评估方法常忽略临床有意义指标。我们系统分析了标注质量、数据集偏差和领域偏移问题。跨数据集评估显示多种模型架构均出现明显外部性能下降,AUPRC与F1分数显著降低。通过源分类模型训练,可近乎完美区分不同数据集,子群分析表明少数群体(如特定年龄、性别)表现更差。两位认证放射科医师的专家评审亦揭示标注与实际诊断存在显著分歧。研究凸显当前基准数据集的临床缺陷,强调需建立经临床验证的数据集和更公平的评估框架。
原文摘要 · Abstract (English)
Artificial intelligence has shown significant promise in chest radiography, where deep learning models can approach radiologist-level diagnostic performance. Progress has been accelerated by large public datasets such as MIMIC-CXR, ChestX-ray14, PadChest, and CheXpert, which provide hundreds of thousands of labelled images with pathology annotations. However, these datasets also present important limitations. Automated label extraction from radiology reports introduces errors, particularly in handling uncertainty and negation, and radiologist review frequently disagrees with assigned labels. In addition, domain shift and population bias restrict model generalisability, while evaluation practices often overlook clinically meaningful measures. We conduct a systematic analysis of these challenges, focusing on label quality, dataset bias, and domain shift. Our cross-dataset domain shift evaluation across multiple model architectures revealed substantial external performance degradation, with pronounced reductions in AUPRC and F1 scores relative to internal testing. To assess dataset bias, we trained a source-classification model that distinguished datasets with near-perfect accuracy, and performed subgroup analyses showing reduced performance for minority age and sex groups. Finally, expert review by two board-certified radiologists identified significant disagreement with public dataset labels. Our findings highlight important clinical weaknesses of current benchmarks and emphasise the need for clinician-validated datasets and fairer evaluation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。