提出新方法提升病理生存模型预测可靠性,可量化个体风险时间范围。
A Multi-Cohort Validation of Censoring-Aware Conformal Lower Predictive Bounds for Pathology Survival Models
- 用冻结的UNI2-h特征+分段生存模型构建置信区间,考虑删失数据影响。
- 在多个癌症队列中,90%以上覆盖率达目标,关键队列表现稳定。
- 适合临床研究者评估患者生存风险,尤其关注删失数据下的可靠性。
全切片生存模型通常仅提供风险排序,缺乏对个体事件时间的校准预测。本文评估了固定阈值drcosarc方法,在五个TCGA队列内部18种配置和三个CPTAC队列外部5种配置下表现。通过区分逆概率删失加权(IPCW)估计与中位下界预测区间(LPB),并基于层次感知的患者集成估计量比较drcosarc与朴素方法的差异。在α=0.1时,KIRC、LUAD和STAD队列中IPCW估计最接近0.90。患者集成的drcosarc-朴素区间在KIRC、KIRP、STAD、UCEC及CPTAC-CCRCC中排除零点,但在内部LUAD、CPTAC-LUAD、CPTAC-UCEC及扩展的内部LUSC中包含零点。在20次重复的低删失半合成数据中,经验覆盖率为0.9129 [0.9053, 0.9207]。探索性分析显示头错误与删失存在交互作用。两队列ABMIL敏感性分析表明,将危险率网格增至K=16后,局部边际IPCW估计超过预设0.87阈值,并产生正向配对LPB差异,但最差组仍低于0.87。总体表现受队列影响,其解释随患者层面单位、估计量及删失假设变化。
原文摘要 · Abstract (English)
Whole-slide survival models commonly provide risk rankings without calibrated statements about individual event times. We evaluate fixed-cutoff drcosarc, a post-hoc conformal wrapper for discrete-time multiple-instance learning survival heads using frozen UNI2-h representations, in an internal 18-configuration sweep across five TCGA cohorts and an external five-configuration evaluation across three CPTAC cohorts. We distinguish configuration--fold--split summaries of the inverse-probability-of-censoring-weighted (IPCW) estimate and median lower predictive bound (LPB) from a hierarchy-aware patient-ensemble estimand of the mean drcosarc--naive LPB difference. At $α=0.1$, the drcosarc IPCW estimate was nearest 0.90 in KIRC, LUAD, and STAD. Patient-ensemble drcosarc--naive intervals excluded zero in KIRC, KIRP, STAD, UCEC, and CPTAC-CCRCC, but included zero in internal LUAD, CPTAC-LUAD, CPTAC-UCEC, and the internal LUSC extension. In a 20-replicate low-censoring semi-synthetic setting with known event times, drcosarc empirical coverage was 0.9129 [0.9053, 0.9207]. An exploratory analysis supported a head-error-by-censoring interaction within that data-generating process. In a two-cohort ABMIL sensitivity analysis, increasing the hazard grid to $K=16$ raised localized marginal IPCW estimates above the prespecified 0.87 threshold and yielded positive paired LPB differences, although worst-group estimates remained below 0.87. Overall, performance was cohort dependent, and its interpretation changed with the patient-level unit, estimand, and censoring assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。