研究胸部X光分类中罕见病患者被漏诊的问题,发现阈值设置影响重大。
Who Gets Missed in the Tail? Thresholded Subgroup Underdiagnosis in Long-Tailed Chest X-ray Classification
- 用分层诊断法分离标签长尾、子组加权和阈值选择的影响
- 在VinDr-CXR上将尾部错误率从0.665降至0.269,性别最差组从0.705降至0.157
- 强调阈值设定对罕见子组公平性的关键作用,适合医疗AI审计人员参考
在胸部X光(CXR)分类中,即使排名性能良好,仍可能遗漏罕见阳性患者,尤其在特定子组内。本文将这一部署前的公平性问题作为审计课题:多标签长尾CXR模型从分数转为决策后,谁会被漏掉?在VinDr-CXR和MIMIC-CXR/CXR-LT数据集上,采用诊断阶梯方法分离类别级长尾损失、子组感知加权、群体鲁棒性与阈值选择的影响。在VinDr-CXR上,先进行群体尾部加权再做尾部感知阈值化,使尾部误报率(FNR)从0.665降至0.269,性别最差组FNR从0.705降至0.157,年龄最差组从0.822降至0.133,同时宏平均AUC(macro-mAP)从0.611升至0.635。在MIMIC-CXR/CXR-LT上,同样方法使尾部FNR从0.866降至0.741,并降低性别、年龄、种族、保险等子组的最差组FNR,但残余漏诊率仍较高。配对自举对比验证了阈值化带来的FNR下降,而GroupDRO基准实验表明仅提升整体群体鲁棒性无法消除稀有子组的漏诊。研究支持一个狭义审计结论:胸部X光中的罕见标签公平性取决于发现类型、子组及操作阈值,而非仅标签频率或排序指标。
原文摘要 · Abstract (English)
In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups. We study this pre-deployment fairness problem as an audit question: after a long-tailed multi-label CXR model is converted from scores into decisions, who is missed? Across VinDr-CXR and MIMIC-CXR/CXR-LT, we use a diagnostic ladder to separate class-level long-tail losses, subgroup-aware weighting, group robustness, and threshold selection. On VinDr-CXR, group-tail weighting followed by tail-aware thresholding reduces tail FNR from 0.665 to 0.269, sex worst-group FNR from 0.705 to 0.157, and age worst-group FNR from 0.822 to 0.133, while macro-mAP increases from 0.611 to 0.635. On MIMIC-CXR/CXR-LT, the same score-to-threshold comparison reduces tail FNR from 0.866 to 0.741 and lowers worst-group FNR across sex, age, race, and insurance; residual missed-positive rates nevertheless remain high. Paired bootstrap contrasts on VinDr support the thresholded FNR reductions, and GroupDRO reference runs indicate that aggregate group robustness alone does not remove rare subgroup misses in this setting. The study supports a narrow audit claim: rare-label fairness in CXR depends jointly on the finding, subgroup, and operating threshold, not on label frequency or ranking metrics alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。