arXiv:2409.12116cs.LGcs.CY2024-09被引 4

强基线模型能更好揭示医疗AI价值,助力临床落地。

Stronger Baseline Models -- A Key Requirement for Aligning Machine Learning Research with Clinical Utility

  • 用更强基线对比,更真实反映新模型优势
  • 弱基线会掩盖实际算法价值,误导研究评估
  • 适合关注临床应用的医疗AI研究者与从业者

近年来,机器学习(ML)研究迅速增长,得益于其在多个领域的预测建模成功。然而,在高风险临床环境中部署ML模型仍面临诸多障碍,包括模型缺乏透明性(难以审计推理过程)、训练数据需求大且数据孤岛化、以及衡量模型实用性的指标复杂等。本文通过一系列案例研究实证表明,在医疗ML评估中引入更强的基线模型,能有效帮助从业者应对这些挑战。我们发现,常见做法是省略基线或仅与弱基线(如未优化的线性模型)比较,这会掩盖研究文献中提出的新方法的实际价值。基于这些发现,本文提出若干最佳实践,以促进从业者更有效地研究和部署医疗领域中的机器学习模型。

原文摘要 · Abstract (English)

Machine Learning (ML) research has increased substantially in recent years, due to the success of predictive modeling across diverse application domains. However, well-known barriers exist when attempting to deploy ML models in high-stakes, clinical settings, including lack of model transparency (or the inability to audit the inference process), large training data requirements with siloed data sources, and complicated metrics for measuring model utility. In this work, we show empirically that including stronger baseline models in healthcare ML evaluations has important downstream effects that aid practitioners in addressing these challenges. Through a series of case studies, we find that the common practice of omitting baselines or comparing against a weak baseline model (e.g. a linear model with no optimization) obscures the value of ML methods proposed in the research literature. Using these insights, we propose some best practices that will enable practitioners to more effectively study and deploy ML models in clinical settings.

医疗AI模型评估基线对比临床落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。