arXiv:2504.16185stat.MLcs.LG2025-04中稿 · publication in the…被引 7

AUC在罕见事件中是否可靠?关键看事件数量而非发生率。

Behavior of prediction performance metrics with rare events

  • 通过模拟研究发现,AUC稳定性取决于事件总数,而非事件发生率。
  • 当事件数达1000时,AUC偏差接近零,表现稳定。
  • 敏感性依赖事件数,特异性依赖非事件数,其他指标受发生率影响。

目标:二元结局预测模型常报告受试者工作特征曲线下面积(AUC)。近期研究指出,在罕见事件场景下AUC可能误导性能评估,而此类场景在临床中常见。本文旨在确定AUC的偏倚与方差是由事件数还是事件率驱动,并考察阳性预测值、准确率、敏感性和特异性等常用指标的表现。研究设计与背景:基于精神健康研究网络数据进行拟合模拟研究,数据包含149个预测变量,关注自杀未遂事件,原始数据中事件率为0.92%。结果:研究表明,AUC在罕见事件中的不良表现——表现为经验偏倚、交叉验证AUC估计的变异性及置信区间经验覆盖率——由事件数决定,而非事件率。敏感性表现依赖事件数,特异性表现依赖非事件数。其他指标如阳性预测值和准确率,即使在大样本中仍受事件率影响。结论:只要事件总数适中,AUC在罕见事件设置中依然可靠;模拟中事件数达1000时,偏差接近零。

原文摘要 · Abstract (English)

Objective: Area under the receiving operator characteristic curve (AUC) is commonly reported alongside prediction models for binary outcomes. Recent articles have raised concerns that AUC might be a misleading measure of prediction performance in the rare event setting. This setting is common since many events of clinical importance are rare. We aimed to determine whether the bias and variance of AUC are driven by the number of events or the event rate. We also investigated the behavior of other commonly used measures of prediction performance, including positive predictive value, accuracy, sensitivity, and specificity. Study Design and Setting: We conducted a simulation study to determine when or whether AUC is unstable in the rare event setting by varying the size of datasets used to train and evaluate prediction models. This plasmode simulation study was based on data from the Mental Health Research Network; the data contained 149 predictors and the outcome of interest, suicide attempt, which had event rate 0.92\% in the original dataset. Results: Our results indicate that poor AUC behavior -- as measured by empirical bias, variability of cross-validated AUC estimates, and empirical coverage of confidence intervals -- is driven by the number of events in a rare-event setting, not event rate. Performance of sensitivity is driven by the number of events, while that of specificity is driven by the number of non-events. Other measures, including positive predictive value and accuracy, depend on the event rate even in large samples. Conclusion: AUC is reliable in the rare event setting provided that the total number of events is moderately large; in our simulations, we observed near zero bias with 1000 events.

AUC罕见事件预测性能模拟研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。