arXiv:2506.02075stat.MEcs.LG2025-06被引 8

警告:生存分析评估中过度依赖C-index可能误导结果

Position: Stop Chasing the C-index when Evaluating Survival Analysis Models

  • 提出评估需匹配建模假设,避免指标与目标错位
  • 实验证明假设不符时模型比较结果会严重失真
  • 适合关注评估可靠性、避免常见陷阱的研究者

当前生存分析的评估方法存在根本性问题:评价指标的使用常与建模目标不一致,且对删失(censoring)的假设往往隐含或缺乏依据。这导致报告性能可能具有误导性,无法回答原始科学问题。本文批判性审视生存分析评估实践,指出删失使评估区别于标准回归或分类。重点分析了广泛使用的协和性度量(如C-index),揭示其在文献中被过度依赖。为此,我们提出一组关键理想特性,并引入双螺旋阶梯模型,强调评估有效性取决于指标与建模假设的一致性。通过受控实验表明,这种一致性缺失会导致模型比较结果失真。最后提供实用指导,帮助研究者选择合适的评估方式。

原文摘要 · Abstract (English)

The current state of evaluation in survival analysis is plagued by the persistent use of evaluation metrics in ways that are misaligned with the stated modeling objective. In addition, many such evaluations are based on censoring assumptions that are left implicit or unjustified. This means that the reported performance can be misleading and may fail to answer the scientific or modeling question the evaluation was intended to address. In this position paper, we critically examine evaluation practices in survival analysis and highlight how censoring makes evaluation fundamentally different from standard regression or classification. We place particular focus on concordance-based measures, such as the C-index, which we show are heavily overused in the literature. To help identify appropriate metrics, we propose a set of key desiderata and introduce a double-helix ladder, in which valid evaluation requires alignment between metric and modeling assumptions. Through controlled experiments, we show that violations of this alignment can lead to misleading model comparisons. We conclude by providing practical guidance on how to evaluate a survival model.

生存分析评估方法模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。