arXiv:2602.08171cs.LGstat.AP2026-02

用因果机器学习区分治疗差异是否真能提升决策效果。

A Causal Machine Learning Framework for Treatment Personalization in Clinical Trials: Application to Ulcerative Colitis

  • 分三步评估:找相关特征、验统计显著性、测实际疗效改善
  • 内镜指标虽与疗效差异相关,但用于选药反而降低治愈率
  • 临床指标(如钙结合蛋白)才是真正影响治疗选择的关键

随机对照试验评估平均治疗效果,但个体反应差异推动个性化治疗。关键问题是:统计上可检测的异质性是否真的能改进治疗决策——二者可能矛盾。本文提出模块化因果机器学习框架,分别回答:通过置换重要性识别预测异质性的特征,用最佳线性预测检验(BLP)评估统计显著性,以双重稳健政策评估衡量基于异质性决策能否改善患者结局。将该框架应用于溃疡性结肠炎的UNIFI维持期试验数据,比较安慰剂、每12周标准剂量乌司奴单抗、每8周强化剂量乌司奴单抗三种方案。使用交叉拟合X-learner模型,输入包括基线人口学、用药史、第8周临床评分、实验室生物标志物及视频衍生内镜特征。BLP检验发现内镜特征与乌司奴单抗相比安慰剂的疗效异质性有强关联;但双重稳健政策评估显示,引入内镜特征并未提升预期缓解率,且外部折叠多臂评估表现更差。诊断性对比显示,内镜评分主要反映疾病严重程度,有助于预测未治疗患者结局,但在治疗选择中引入噪声;而临床变量(粪便钙结合蛋白、年龄、CRP)捕捉了真正与决策相关的变异。结果表明,临床试验中的因果机器学习应包含政策层面评估,而不仅是异质性检验。

原文摘要 · Abstract (English)

Randomized controlled trials estimate average treatment effects, but treatment response heterogeneity motivates personalized approaches. A critical question is whether statistically detectable heterogeneity translates into improved treatment decisions -- these are distinct questions that can yield contradictory answers. We present a modular causal machine learning framework that evaluates each question separately: permutation importance identifies which features predict heterogeneity, best linear predictor (BLP) testing assesses statistical significance, and doubly robust policy evaluation measures whether acting on the heterogeneity improves patient outcomes. We apply this framework to patient-level data from the UNIFI maintenance trial of ustekinumab in ulcerative colitis, comparing placebo, standard-dose ustekinumab every 12 weeks, and dose-intensified ustekinumab every 8 weeks, using cross-fitted X-learner models with baseline demographics, medication history, week-8 clinical scores, laboratory biomarkers, and video-derived endoscopic features. BLP testing identified strong associations between endoscopic features and treatment effect heterogeneity for ustekinumab versus placebo, yet doubly robust policy evaluation showed no improvement in expected remission from incorporating endoscopic features, and out-of-fold multi-arm evaluation showed worse performance. Diagnostic comparison of prognostic contribution against policy value revealed that endoscopic scores behaved as disease severity markers -- improving outcome prediction in untreated patients but adding noise to treatment selection -- while clinical variables (fecal calprotectin, age, CRP) captured the decision-relevant variation. These results demonstrate that causal machine learning applications to clinical trials should include policy-level evaluation alongside heterogeneity testing.

因果推断个性化治疗临床试验机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。