arXiv:2607.10466cs.LGstat.AP2026-07

医学生存模型易受时间截断泄漏影响,导致预测偏差。

Pitfalls of Administrative Censoring in Survival Models with Time-Indexed Inputs

论文配图:Pitfalls of Administrative Censoring in Survival Models with Time-Indexed Inputs
图 1 · 摘自论文原文
  • 利用临床数据时间线索误判风险,而非真实事件概率。
  • 模拟与真实数据均显示截断泄漏会夸大AUC和C指数。
  • 建议数据至少提供预测期两倍的随访时间以避免偏差。

生存模型可基于部分观测数据建模事件发生时间,广泛应用于癌症风险、疾病进展、治疗反应和死亡率等临床预测。近期模型常使用特定就诊时采集的丰富输入,如医学影像、检验结果、电子病历快照或传感器数据。在大型回顾性数据集中,这些输入跨越多年,可能隐含采集时间信息(如设备、流程、患者构成或临床实践变化)。当结局仅观察至固定研究截止日期时,较新记录必然具有更短潜在随访期,导致模型可能通过输入推断记录时间,从而学习随访时长而非真实风险。我们称此为行政截断泄漏。本文界定其发生条件,区分于经典信息性删失和真实风险变化,并提出检测方法。模拟显示该泄漏会抬高固定时限AUC,且在现实随访模式下影响哈雷尔C指数;真实乳腺钼靶队列亦呈现相同现象。研究建议:对n年预测任务,数据集应至少提供截至最新输入日期后n年的潜在随访时间,否则模型可能受行政截断泄漏引起的偏差影响。

原文摘要 · Abstract (English)

Survival models can model time-to-event outcomes using partially observed data. They are widely used in clinical prediction, including cancer risk, disease progression, treatment response, and mortality. Recent models often rely on rich inputs collected at a specific clinical encounter, such as medical images, laboratory tests, electronic health record snapshots, or sensor measurements. In large retrospective datasets, these inputs are usually collected over many calendar years. As a result, they may contain clues about when they were acquired through changes in devices, protocols, documentation, patient mix, or clinical practice. This creates a potential failure mode when outcomes are observed only up to a fixed study end date. More recent records necessarily have less potential follow-up than older records. A model that can infer the record date from the input may therefore learn to predict how much follow-up was available rather than the patient's true risk of experiencing the event. We call this failure mode administrative-cutoff leakage. In this paper, we characterize when this leakage can occur, distinguish it from classical informative censoring and genuine temporal changes in risk, and propose practical ways to detect it. In simulations, we show that administrative-cutoff leakage can inflate fixed-horizon AUC and can also affect Harrell's C-index under realistic follow-up patterns. We then demonstrate the same behavior in a real mammography cohort. These results motivate a simple design principle for survival prediction: for an n-year prediction task, the dataset should provide at least n years of potential follow-up after the latest input date. Otherwise, the models may be subject to bias induced by administrative-cutoff leakage.

生存分析临床预测数据偏差随访设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。