arXiv:2505.04720cs.CV2025-05被引 9

80%的医学影像AI论文夸大性能优势,真实效果可能只是运气。

False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims

  • 用贝叶斯方法评估模型真实排名是否偶然出现
  • 86%分类论文和53%分割论文存在超5%的假性优势风险
  • 提醒研究者警惕仅靠平均性能就宣称超越对手的做法

在医学影像AI研究中,性能比较常以相对提升作为优越性依据,但多数结论仅基于经验均值。本文分析了代表性医学影像论文集,采用贝叶斯方法结合报告结果与实测模型一致性,评估新方法真实优于当前最优的可能性。结果显示,超过80%的论文声称性能提升;其中86%的分类论文和53%的分割论文存在超过5%的虚假优势概率。这揭示了当前基准测试的重大缺陷:多数性能领先宣称缺乏可靠证据,可能误导后续研究方向。

原文摘要 · Abstract (English)

Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common performance metrics. However, such claims frequently rely solely on empirical mean performance. In this paper, we investigate whether newly proposed methods genuinely outperform the state of the art by analyzing a representative cohort of medical imaging papers. We quantify the probability of false claims based on a Bayesian approach that leverages reported results alongside empirically estimated model congruence to estimate whether the relative ranking of methods is likely to have occurred by chance. According to our results, the majority (>80%) of papers claims outperformance when introducing a new method. Our analysis further revealed a high probability (>5%) of false outperformance claims in 86% of classification papers and 53% of segmentation papers. These findings highlight a critical flaw in current benchmarking practices: claims of outperformance in medical imaging AI are frequently unsubstantiated, posing a risk of misdirecting future research efforts.

医学影像性能验证假阳性贝叶斯分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。