跨国家儿科X光肺炎模型评估,发现性能衰减与决策行为变化
Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery
- 分项测试模型在不同国家数据上的判别力、概率校准和决策点稳定性
- 跨数据集性能下降显著,冻结阈值下敏感度降至6.2%甚至0%
- 少量标注可恢复敏感度但伴随特异度波动,需警惕误报风险
背景与目标:医学影像AI的外部评估常被简化为判别能力。本文提出一种计算协议,分别评估儿科肺炎分类模型在三个国家数据集上的判别力、概率校准、固定决策点迁移、捷径信号关联性及小样本恢复能力。方法:剔除完全重复图像后,5,824张广州影像用于源模型开发与内部测试。冻结的三种子随机数的DenseNet121双视角集成模型零样本评估于孟加拉国的BDCXR-3257(n=3,257)和越南的整合版VinDr-PCXR/PediCXR测试集(n=1,077)。匹配种子42的变体测试架构鲁棒性。对BDCXR的二次分析使用固定651图像适应池和2,606图像保留集;163、326、651标签分别代表完整BDCXR数据的5%、10%、20%。结果:内部AUROC达0.976,敏感度95.1%。在BDCXR和VinDr-PCXR上,AUROC分别为0.798和0.742,而冻结阈值下的敏感度分别降至6.2%和0%。从源数据到BDCXR,全图基线(0.961→0.749)、无门控双视图模型(0.977→0.766)和门控MixStyle模型(0.966→0.789)均出现显著性能下降。仅用163个标签进行Platt校准可维持AUROC并使保留集敏感度提升至88.3%,但特异度仅47.9%,警报率高达78.5%。200次重复163标签拟合证实敏感度可恢复,但特异度波动大。结论:跨国数据分布差异对排序、概率对齐及源定义决策行为的影响各异。迁移研究应分项评估,并量化看似恢复背后的运营负担。
原文摘要 · Abstract (English)
Background and Objective: External evaluation of medical-imaging AI is often collapsed into discrimination. We evaluated a computational protocol that separately tests discrimination, probability calibration, fixed operatingpoint transport, shortcut-associated signal, and limited-label recoverability for pediatric pneumonia classification across datasets from three countries. Methods: After exact-duplicate removal, 5,824 Guangzhou radiographs supported leakage-controlled source development and internal testing. A frozen three-seed DenseNet121 dual-view ensemble was evaluated zero-shot on BDCXR-3257 from Bangladesh (n = 3, 257) and an untouched harmonized VinDr-PCXR/PediCXR test cohort from Vietnam (n = 1, 077). Matched seed-42 variants tested architectural robustness. Secondary BDCXR analyses used a fixed 651-image adaptation pool and 2,606-image hold-out; 163, 326, and 651 labels represented 5%, 10%, and 20% of complete BDCXR. Results: Internal AUROC was 0.976 with 95.1% sensitivity. BDCXR and VinDr-PCXR AUROC were 0.798 and 0.742, while frozen-threshold sensitivity fell to 6.2% and 0%. Source-to-BDCXR AUROC degradation occurred for a full-image baseline (0.961 to 0.749), ungated dual-view model (0.977 to 0.766), and gated MixStyle model (0.966 to 0.789). With 163 BDCXR labels, Platt recalibration preserved AUROC while increasing held-out sensitivity to 88.3%, but specificity was 47.9% and the alert rate was 78.5%. Two hundred repeated 163-label fits confirmed sensitivity recovery but substantial specificity variability. Conclusions: Cross-dataset shifts across countries affected ranking, probability alignment, and source-defined decision behavior differently. Transport studies should evaluate these components separately and quantify the operational burden of apparent recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。