PISA构建可解释的生存分析流水线,兼顾模型精度与临床可读性。
PISA: An AI Pipeline for Interpretable-by-design Survival Analysis Providing Multiple Complexity-Accuracy Trade-off Models
- 通过多特征多目标工程生成多种复杂度-精度权衡的生存模型
- 在两个临床数据集上达到顶尖性能并输出直观分层流程图
- 适合需要可解释性临床决策支持的研究者和医生
生存分析在临床研究中至关重要,可辅助患者预后判断、治疗决策与资源分配。准确的时间到事件预测不仅能提升生活质量,还能揭示影响临床实践的风险因素。为使模型在医疗领域具有实际价值,可解释性尤为关键:预测结果需追溯至患者个体特征,风险因素应可识别以提供可操作的洞见。传统生存模型难以捕捉非线性交互,而现代深度学习方法虽强大却缺乏可解释性。我们提出一种可解释性设计的生存分析流水线(PISA),能生成多种复杂度与性能权衡的生存模型。通过多特征、多目标特征工程,将患者特征与生存数据转化为多个模型,并将每个模型转换为基于Kaplan-Meier曲线的简单患者分层流程图,且不牺牲性能。尽管PISA对模型无偏好,我们通过Cox回归与浅层生存树的应用展示了其灵活性,后者避免了比例风险假设。在两个临床基准数据集上的应用表明,PISA生成了可解释模型与直观分层图,同时达到最先进性能。进一步回溯一项既往科室研究,验证了其在真实临床研究中自动化生存分析工作流的能力。
原文摘要 · Abstract (English)
Survival analysis is central to clinical research, informing patient prognoses, guiding treatment decisions, and optimising resource allocation. Accurate time-to-event predictions not only improve quality of life but also reveal risk factors that shape clinical practice. For these models to be relevant in healthcare, interpretability is critical: predictions must be traceable to patient-specific characteristics, and risk factors should be identifiable to generate actionable insights for both clinicians and researchers. Traditional survival models often fail to capture non-linear interactions, while modern deep learning approaches, though powerful, are limited by poor interpretability. We propose a Pipeline for Interpretable Survival Analysis (PISA) - a pipeline that provides multiple survival analysis models that trade off complexity and performance. Using multiple-feature, multi-objective feature engineering, PISA transforms patient characteristics and time-to-event data into multiple survival analysis models, providing valuable insights into the survival prediction task. Crucially, every model is converted into simple patient stratification flowcharts supported by Kaplan-Meier curves, whilst not compromising on performance. While PISA is model-agnostic, we illustrate its flexibility through applications of Cox regression and shallow survival trees, the latter avoiding proportional hazards assumptions. Applied to two clinical benchmark datasets, PISA produced interpretable survival models and intuitive stratification flowcharts whilst achieving state-of-the-art performances. Revisiting a prior departmental study further demonstrated its capacity to automate survival analysis workflows in real-world clinical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。