arXiv:2604.20296stat.MLcs.LG2026-04

将Cox模型引入在线学习,实现动态治疗决策优化。

Online Survival Analysis: A Bandit Approach under Cox PH Model

  • 基于Cox PH模型设计在线学习算法,处理延迟反馈与删失数据。
  • 三种经典强化学习算法适配后,实现亚线性后悔界,学习效果稳定。
  • 适合医疗领域动态干预策略研究,可快速收敛至近优治疗方案。

生存分析是建模右删失条件下时间-事件数据的常用统计框架。经典方法如Cox比例风险(Cox PH)模型提供了半参数方式估计协变量对风险函数的影响。尽管重要,生存分析在在线学习场景中仍基本未被探索,尤其是在需随新数据到来顺序决策以优化治疗的带宽(bandit)框架下。本文首次尝试在纯在线学习设定下整合生存分析与Cox PH模型,解决包括异步入组、延迟反馈和右删失在内的关键挑战。我们适配了三种经典的带宽算法以平衡探索与利用,并给出亚线性后悔边界理论保证。通过大量模拟及使用SEER癌症数据的半真实实验表明,该方法能快速有效学习接近最优的治疗策略。

原文摘要 · Abstract (English)

Survival analysis is a widely used statistical framework for modeling time-to-event data under censoring. Classical methods, such as the Cox proportional hazards (Cox PH) model, offer a semiparametric approach to estimating the effects of covariates on the hazard function. Despite its importance, survival analysis has been largely unexplored in online settings, particularly within the bandit framework, where decisions must be made sequentially to optimize treatments as new data arrive over time. In this work, we take an initial step toward integrating survival analysis into a purely online learning setting under the Cox PH model, addressing key challenges including staggered entry, delayed feedback, and right censoring. We adapt three canonical bandit algorithms to balance exploration and exploitation, with theoretical guarantees of sublinear regret bounds. Extensive simulations and semi-real experiments using SEER cancer data demonstrate that our approach enables rapid and effective learning of near-optimal treatment policies.

生存分析在线学习带宽算法Cox模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。