arXiv:2412.10288cs.LGstat.ME2024-12综述被引 32

为医疗预测模型选对评估指标,避免误判伤人。

Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

  • 区分统计与临床决策性能,选对评估指标
  • 17个指标兼具正确概率与成本考量优势
  • 推荐报告AUC、校准图、净收益等关键指标

大量用于评估预测型人工智能(AI)模型性能的指标已被提出。在医疗实践中应用预测模型时,选择恰当的性能指标至关重要,因表现不佳的模型可能伤害患者并增加成本。本文旨在评估经典与现代性能指标在医疗预测模型验证中的优劣,聚焦二分类任务。讨论了涵盖五类性能域(区分度、校准、整体、分类、临床效用)的32个指标及其配套图形评估方法。前四类反映统计性能,第五类反映决策分析性能。强调选择指标时两个关键特征:(1)其期望值是否在使用真实概率时最优(即‘合适’指标);(2)是否同时考虑误分类成本。其中17个指标具备双重特性,14个具备单一特性,仅F1指标两者皆无。所有分类指标(如准确率、F1)在非0.5或先验概率的临床阈值下均不适用。建议必须报告:AUROC、校准图、以决策曲线分析为基础的净收益,以及按结果类别划分的概率分布图。

原文摘要 · Abstract (English)

A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be used in medical practice, because poorly performing models may harm patients and lead to increased costs. We aim to assess the merits of classic and contemporary performance measures when validating predictive AI models for use in medical practice. We focus on models with a binary outcome. We discuss 32 performance measures covering five performance domains (discrimination, calibration, overall, classification, and clinical utility) along with accompanying graphical assessments. The first four domains cover statistical performance, the fifth domain covers decision-analytic performance. We explain why two key characteristics are important when selecting which performance measures to assess: (1) whether the measure's expected value is optimized when it is calculated using the correct probabilities (i.e., a "proper" measure), and (2) whether they reflect either purely statistical performance or decision-analytic performance by properly considering misclassification costs. Seventeen measures exhibit both characteristics, fourteen measures exhibited one characteristic, and one measure possessed neither characteristic (the F1 measure). All classification measures (such as classification accuracy and F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence. We recommend the following measures and plots as essential to report: AUROC, calibration plot, a clinical utility measure such as net benefit with decision curve analysis, and a plot with probability distributions per outcome category.

医疗AI模型评估临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。