提出两种方法提升预测模型在不同人群间的适用性。
Robustifying and Selecting Cohort-Appropriate Prognostic Models under Distributional Shifts

- 以元分析分布为基准优化模型,增强跨人群泛化能力。
- 发现分布差异越大,校准性能越差,且与临床效用相关。
- 提供简单指标帮助医生选最适合目标人群的已有模型。
外部验证被广泛视为预后模型评估的金标准。本研究挑战了‘成功外部校准即代表模型泛化性好’这一假设,提出两种互补策略以提升预后模型在不同队列间的可迁移性。基于六组来自三甲医院的真实外科队列数据,我们检验了外部校准是否依赖于训练与验证队列间协变量和结局的相似性,采用KL散度量化分布差异,以整合校准指数(ICI)评估校准效果。从模型开发者角度,通过调整模型以逼近元分析推导出的协变量与结局分布,构建‘平均最优’预后模型。从使用者角度,提出一种简单的队列结局相似性度量,用于从已发表模型中筛选最适配当前队列的模型,兼顾校准与临床实用性。结果显示,分布失配程度越高,校准性能越差;在仅手术组(Spearman ρ=0.614, p=0.004)和联合治疗组(Spearman ρ=0.738, p<0.001)中,KL散度与ICI显著正相关。基于元分析加权的方法在多数场景下改善了校准效果,且未明显影响判别能力,在聚合外部人群中表现尤为显著(p=0.037)。在协变量与结局更相似的队列中开发的模型,其校准误差更低(手术组ρ=0.803, p<0.001;联合治疗组ρ=0.737, p<0.001),并在决策曲线分析(DCA)中展现更高临床效用。
原文摘要 · Abstract (English)
External validation is widely regarded as the gold standard for prognostic model evaluation. In this study, we challenge the assumption that successful external calibration guarantees model generalizability and propose two complementary strategies to improve transportability of prognostic models across cohorts. Using six real-world surgical cohorts from tertiary academic centers, we tested whether successful external calibration depends largely on similarity in covariates and outcomes between training and validation cohorts, quantified using Kullback-Leibler (KL) divergence, with calibration assessed by the Integrated Calibration Index (ICI). From the model-developer's perspective, we trained the "best-on-average" prognostic model by tuning toward a meta-analysis-derived covariate and outcome distribution as an approximation of the broader target population. From the end-user perspective, we proposed a simple measure for cohort outcome similarity to identify, among published models, the one most suitable for a given target cohort in terms of both calibration and clinical utility. External calibration worsened as distributional mismatch increased. Higher KL divergence was associated with higher ICI in both surgery-alone (Spearman $ρ=0.614$, $p=0.004$) and surgery + adjuvant chemotherapy cohorts (Spearman $ρ=0.738$, $p<0.001$). Meta-analysis-informed weighting improved calibration in most settings without materially affecting discrimination, with the clearest benefit when evaluated on the aggregated external population ($p=0.037$). Models developed in more similar cohorts achieved lower ICI in surgery-alone (Spearman $ρ=0.803$, $p<0.001$) and surgery + adjuvant chemotherapy cohorts (Spearman $ρ=0.737$, $p<0.001$), and provided greater clinical utility on DCA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。