医学AI竞赛排名需改进:统计显著性、指标合理性与公平性
Auditing Significance, Metric Choice, and Demographic Fairness in Medical AI Challenges
- 用成对显著性分析揭示模型性能差异的可靠性
- 换用器官适配指标后,前四名模型排名逆转
- 发现超过一半的MONAI模型存在性别种族不公平
开放挑战赛已成为医疗AI方法比较的默认标准。然而,医疗AI排行榜存在三大持续问题:(1) 分数差距很少经过统计显著性检验,排名稳定性未知;(2) 对所有器官使用单一平均指标,掩盖了临床重要的边界误差;(3) 缺乏对交叉人口学特征的表现报告,隐藏了公平性与公正性差距。我们提出RankInsight,一个开源工具包以解决这些问题。RankInsight (1) 计算成对显著性图,显示nnU-Net系列在高统计置信度下优于视觉语言和MONAI提交模型;(2) 使用器官适配指标重新计算排行榜,当用NSD替代Dice评估管状结构时,前四名模型顺序发生反转;(3) 审计交叉公平性,在我们自有的约翰霍普金斯医院数据集上发现,超过一半的MONAI相关条目存在最大性别-种族差异。RankInsight工具包已公开发布,可直接应用于过往、正在进行及未来的挑战赛,使组织者和参与者能发布具有统计严谨性、临床意义和人口公平性的排名。
原文摘要 · Abstract (English)
Open challenges have become the de facto standard for comparative ranking of medical AI methods. Despite their importance, medical AI leaderboards exhibit three persistent limitations: (1) score gaps are rarely tested for statistical significance, so rank stability is unknown; (2) single averaged metrics are applied to every organ, hiding clinically important boundary errors; (3) performance across intersecting demographics is seldom reported, masking fairness and equity gaps. We introduce RankInsight, an open-source toolkit that seeks to address these limitations. RankInsight (1) computes pair-wise significance maps that show the nnU-Net family outperforms Vision-Language and MONAI submissions with high statistical certainty; (2) recomputes leaderboards with organ-appropriate metrics, reversing the order of the top four models when Dice is replaced by NSD for tubular structures; and (3) audits intersectional fairness, revealing that more than half of the MONAI-based entries have the largest gender-race discrepancy on our proprietary Johns Hopkins Hospital dataset. The RankInsight toolkit is publicly released and can be directly applied to past, ongoing, and future challenges. It enables organizers and participants to publish rankings that are statistically sound, clinically meaningful, and demographically fair.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。