提出新方法发现生存分析中表现优异的子群体,提升模型可解释性。
Subgroup Discovery with the Cox Model
- 引入预测熵和条件秩统计量,改进子群体评估指标
- 在真实与合成数据上验证,能准确识别出真实子群体
- 适合医学、工程等需可解释生存分析的场景
本文研究生存分析中的子群体发现问题,目标是找到一个可解释的数据子集,在其上Cox模型具有高精度。这是首个针对Cox模型子群体发现的研究。现有质量函数无法有效解决该问题。为此,提出两个创新:预期预测熵(EPE),用于评估预测危险函数的生存模型;条件秩统计量(CRS),量化个体点相对于子群体生存时间分布的偏离程度。理论分析表明,二者可克服原有度量的缺陷。提出共八种算法,主算法结合EPE与CRS,可在理想设定下提供理论正确性保证。在合成与真实数据上的实验验证了理论结果:在理想情况下可恢复真实子群体,在实际中相比全数据拟合的Cox模型有更好拟合效果。最后在NASA喷气发动机仿真数据上进行案例研究,发现的子群体揭示了数据中已知的非线性/同质性特征,其结论与实际设计选择一致。
原文摘要 · Abstract (English)
We study the problem of subgroup discovery for survival analysis, where the goal is to find an interpretable subset of the data on which a Cox model is highly accurate. Our work is the first to study this particular subgroup problem, for which we make several contributions. Subgroup discovery methods generally require a "quality function" in order to sift through and select the most advantageous subgroups. We first examine why existing natural choices for quality functions are insufficient to solve the subgroup discovery problem for the Cox model. To address the shortcomings of existing metrics, we introduce two technical innovations: the *expected prediction entropy (EPE)*, a novel metric for evaluating survival models which predict a hazard function; and the *conditional rank statistics (CRS)*, a statistical object which quantifies the deviation of an individual point to the distribution of survival times in an existing subgroup. We study the EPE and CRS theoretically and show that they can solve many of the problems with existing metrics. We introduce a total of eight algorithms for the Cox subgroup discovery problem. The main algorithm is able to take advantage of both the EPE and the CRS, allowing us to give theoretical correctness results for this algorithm in a well-specified setting. We evaluate all of the proposed methods empirically on both synthetic and real data. The experiments confirm our theory, showing that our contributions allow for the recovery of a ground-truth subgroup in well-specified cases, as well as leading to better model fit compared to naively fitting the Cox model to the whole dataset in practical settings. Lastly, we conduct a case study on jet engine simulation data from NASA. The discovered subgroups uncover known nonlinearities/homogeneity in the data, and which suggest design choices which have been mirrored in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。