arXiv:2511.02453cs.LGeess.IV2025-11

提出新统计框架,解决医疗影像模型评估中因随机初始化导致的误判问题。

Accounting for Underspecification in Statistical Claims of Model Superiority

  • 引入训练不确定性作为额外方差分量,改进模型优劣判断的统计框架。
  • 模拟显示种子差异仅1%就显著增加证明性能优越所需的证据强度。
  • 适用于医疗影像领域需严谨验证模型性能的研究者与审稿人。

机器学习方法在医学影像中的应用日益广泛,但许多报告的性能提升缺乏统计稳健性:近期研究指出,微小但显著的性能增益极可能是假阳性。然而,这些分析未考虑模型的‘未指定性’——即表现相似的模型在未见数据上可能因随机初始化或训练动态而行为不同。本文扩展了近期的统计框架,将未指定性作为额外方差成分纳入模型。模拟结果表明,即使种子变异仅约1%,也会显著增加支持性能优越性声明所需证据强度。研究强调,在验证医学影像系统时必须明确建模训练过程的变异性。

原文摘要 · Abstract (English)

Machine learning methods are increasingly applied in medical imaging, yet many reported improvements lack statistical robustness: recent works have highlighted that small but significant performance gains are highly likely to be false positives. However, these analyses do not take \emph{underspecification} into account -- the fact that models achieving similar validation scores may behave differently on unseen data due to random initialization or training dynamics. Here, we extend a recent statistical framework modeling false outperformance claims to include underspecification as an additional variance component. Our simulations demonstrate that even modest seed variability ($\sim1\%$) substantially increases the evidence required to support superiority claims. Our findings underscore the need for explicit modeling of training variance when validating medical imaging systems.

统计评估医疗影像模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。