高区分度的生存模型可能严重失准,真实数据验证了这一潜在误导。
The C-index illusion: discrimination without calibration in published survival models
- 用真实数据复现三类生存模型,检验仅靠区分度评估的可靠性
- 模型区分度接近文献报告值(C=0.9595),但校准性在p<0.001下失败
- 适用于模型审计、风险评估系统验证,尤其关注误判风险的场景
近期研究基于模拟数据提出,仅以区分度(如C-index)评价生存模型会因忽略校准性和时变准确性而产生系统性误导。本文在三个结构不同的真实领域——硬盘故障预测、点对点信用违约、数字平台用户流失——中复现了三项已发表的生存机器学习模型。通过与原始论文的合成实验对比验证评估工具,并在霍尔姆校正后的家族误差率下测试五个预注册假设。其中三项被拒绝(一项临界通过)。一个复现文献区分度(C=0.9595 vs. 0.958)的模型在校准性测试中显著失败(p<0.001);特征消融分析显示无单一特征主导其区分度,表明失准非简单捷径所致。当贷款提前还款被错误视为非信息性删失而非竞争风险时,借款人违约风险估计平均偏高约2个百分点,最危险群体达近4个百分点。平台流失模型的概率估计随预测周期延长而恶化,尽管全局区分度仍在预注册区间内。直接检验指标选择是否反转模型偏好未被拒绝,但统计功效有限(每领域仅2-3个模型)。我们记录的失败模式更应归为对选定模型的过度自信,而非选错模型。论文发布完整预注册评估框架,含代码与注释笔记本,供独立验证与扩展审计。
原文摘要 · Abstract (English)
Recent work has argued normatively, on synthetic data, that evaluating survival models by discrimination alone (concordance index) yields systematically misleading model comparisons, because the metric ignores calibration and time-dependent accuracy. Whether this matters for real, published, non-clinical models has not been tested. We reproduce three published survival-ML models across three structurally distinct domains -- hard-drive failure prediction, peer-to-peer credit default, and user disengagement on digital platforms -- validate our instrument against the anchor paper's own synthetic experiment, and test five pre-registered hypotheses under a Holm-corrected family-wise error rate. Three of five reject (though one pre-registered threshold clears by a narrow margin). A model reproducing the published literature's discrimination almost exactly (C = 0.9595 vs. 0.958 reported) fails a formal calibration test at p < 0.001; a broad feature-ablation search finds no single attribute responsible for its discrimination, so the calibration failure is not a trivial shortcut artifact. A lender's estimated default risk is biased upward by roughly two percentage points, growing to nearly four in the riskiest segment, when loan prepayment is treated as non-informative censoring rather than a competing risk. A platform's churn model shows probability estimates that degrade with the horizon even as global discrimination stays within the pre-registered C-index band. A direct test of whether metric choice inverts model preference does not reject, though with limited power given two to three models per domain; the failure mode we document is better characterized as misplaced confidence in a chosen model than as choosing the wrong one. We release a pre-registered evaluation harness with full code and an annotated notebook, so these results can be verified independently and the audit extended.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。