arXiv:2507.17686stat.MLcs.LG2025-07被引 1

用机器学习修正风险集变化,让生存分析中的危险比更可信。

Debiased maximum-likelihood estimators for hazard ratios under kernel-based machine-learning adjustment

  • 抛弃传统基线风险,用核方法建模治疗导致的风险集动态变化。
  • 提出无偏最大似然估计器,模拟验证其偏差极小,结果准确。
  • 适合处理真实世界观察数据中的动态治疗与复杂协变量问题。

以往研究指出,基于Cox模型估算的治疗组间危险比不可靠,因其未指定的基线风险无法识别治疗分配和未观测因素引起的风险集组成时变。为缓解此问题,尤其在具有未控动态治疗和实时测量多协变量的观察性研究中,本文摒弃基线风险,采用基于核的机器学习显式建模有或无潜变量时的风险集变化。在厘清危险比可因果解释的前提后,基于奈曼正交性构建了无偏最大似然估计器,并证明了必要的收敛性。数值模拟表明该方法能以最小偏差识别真实危险比。这些结果为现代流行病学中处理未控观察数据的因果推断提供了新工具。

原文摘要 · Abstract (English)

Previous studies have shown that hazard ratios between treatment groups estimated with the Cox model are uninterpretable because the unspecified baseline hazard of the model fails to identify temporal change in the risk set composition due to treatment assignment and unobserved factors among multiple, contradictory scenarios. To alleviate this problem, especially in studies based on observational data with uncontrolled dynamic treatment and real-time measurement of many covariates, we propose abandoning the baseline hazard and using kernel-based machine learning to explicitly model the change in the risk set with or without latent variables. For this framework, we clarify the context in which hazard ratios can be causally interpreted, and then develop a method based on Neyman orthogonality to compute debiased maximum-likelihood estimators of hazard ratios, proving necessary convergence results. Numerical simulations confirm that the proposed method identifies the true hazard ratios with minimal bias. These results lay the foundation for developing a useful, alternative method for causal inference with uncontrolled, observational data in modern epidemiology.

生存分析因果推断机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。