arXiv:2607.26333eess.IVcs.CV2026-07

评估标准不同,模型表现排名可能完全颠倒,影响临床应用选择。

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

论文配图:Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance
图 1 · 摘自论文原文
  • 用专家图像和报告标注对比,验证标签来源对模型评价的影响
  • 常用图像质量指标如SSIM、PSNR与医生判断常不一致
  • 评估参考标准应作为临床有效性核心,需明确说明依据

胸部X光(CXR)机器学习依赖自动化评估,通常使用报告生成的标签或通用图像质量指标来近似临床判断。然而,这些指标可能无法可靠反映真实临床决策。本文系统研究了评估参考标准的选择如何影响病理分类与图像质量评估(IQA)中的模型性能与排名。在剑桥大学医院临床队列中,我们收集了配对的专家图像与报告标注,并对MIMIC-CXR数据集子集及诊断图像质量进行了专家评分。结果显示,对于监督式图像分类器(ResNet、DenseNet)及多种零样本与微调的视觉-语言模型(MedKLIP、GLoRIA、ConVIRT),改变标签来源不仅显著改变性能估计,也导致模型排名变化。同时,图像质量评估指标与专家判断的对齐程度高度依赖于具体度量方式,常用指标如SSIM和PSNR常无法匹配医生对诊断可用性的判断。结果表明,评估选择至关重要:它决定了哪些模型被视为最优,进而影响后续开发或部署。因此,评估参考标准应作为胸部X光机器学习临床有效性的核心组成部分,需根据病种、任务和下游临床用途合理论证。

原文摘要 · Abstract (English)

Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.

医学影像评估标准临床相关性模型排名

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。