arXiv:2603.00192cs.LGstat.AP2026-03

提出新评估框架,检测医疗机器学习中个体预测的不稳定性。

Diagnostics for Individual-Level Prediction Instability in Machine Learning for Healthcare

  • 用ePIW和eDFR量化同一患者不同训练下的风险估计波动。
  • 神经网络个体预测不稳定程度可达重采样数据集的水平。
  • 适用于关注临床决策可靠性的医疗AI开发者与研究者。

在医疗领域,预测模型日益用于个体患者决策,但对其风险估计变异性的关注不足。对于当前标准的高参数化机器学习模型,优化与初始化引入的随机性可导致相同患者的预测结果显著不同,而这一问题常被忽略。现有评估依赖聚合指标(如对数损失、准确率),无法捕捉个体层面的不稳定性。我们提出评估框架,采用两个互补诊断:经验预测区间宽度(ePIW)衡量连续风险估计的波动,经验决策翻转率(eDFR)测量基于阈值的临床决策变化。在模拟数据和GUSTO-I临床数据集上应用发现,仅因优化与初始化引起的随机性,即可产生与重采样训练集相当的个体预测差异。神经网络在个体风险预测上的不稳定性远高于逻辑回归。靠近临床决策阈值的风险估计不稳定性会改变治疗建议。因此,稳定性诊断应纳入常规模型验证,以评估临床可靠性。

原文摘要 · Abstract (English)

In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk estimates and its impact on treatment decisions. For overparameterized models, now standard in machine learning, a substantial source of variability often goes undetected. Even when the data and model architecture are held fixed, randomness introduced by optimization and initialization can lead to materially different risk estimates for the same patient. This problem is largely obscured by standard evaluation practices, which rely on aggregate performance metrics (e.g., log-loss, accuracy) that are agnostic to individual-level stability. As a result, models with indistinguishable aggregate performance can nonetheless exhibit substantial procedural arbitrariness, which can undermine clinical trust. We propose an evaluation framework that quantifies individual-level prediction instability by using two complementary diagnostics: empirical prediction interval width (ePIW), which captures variability in continuous risk estimates, and empirical decision flip rate (eDFR), which measures instability in threshold-based clinical decisions. We apply these diagnostics to simulated data and GUSTO-I clinical dataset. Across observed settings, we find that for flexible machine-learning models, randomness arising solely from optimization and initialization can induce individual-level variability comparable to that produced by resampling the entire training dataset. Neural networks exhibit substantially greater instability in individual risk predictions compared to logistic regression models. Risk estimate instability near clinically relevant decision thresholds can alter treatment recommendations. These findings that stability diagnostics should be incorporated into routine model validation for assessing clinical reliability.

医疗AI模型稳定风险预测诊断工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。