arXiv:2604.10412stat.MLcs.LG2026-04

提出可有效估计治疗效果异质性的新方法,适用于风险比与比值比。

Orthogonal machine learning for conditional odds and risk ratios

  • 基于双重稳健与正交风险函数,扩展到比值比和风险比的估计
  • 在复杂数据下显著降低偏差与均方误差,优于传统参数模型
  • 适合关注精准医疗中个体化治疗决策的研究者

条件效应常用于理解治疗效果在不同群体间的差异,是精准干预的关键。本文系统回顾现有方法并提出针对比值比(OR)和风险比(RR)的新估计方法。尽管条件平均治疗效应(ATE)已有广泛研究,但其在OR与RR上的最新技术如双重稳健变换或正交风险函数尚未被推广。本文首次将这些方法拓展至OR与RR,推导出相应的正交风险函数,并证明其伪结果满足类似ATE的二阶条件均值余项性质。通过涵盖数百种数据生成分布的非参数蒙特卡洛模拟,系统比较了条件ATE、OR与RR的多种估计器。结果显示,在真实世界复杂场景中,所提非参数方法显著降低偏差与均方误差,远超传统参数模型。基于美国国家健康与营养调查(NHANES)数据,分析体力活动与睡眠困扰的关系,发现新方法揭示了传统回归忽略的显著治疗异质性,优化了治疗决策规则,凸显数据自适应方法对精准健康研究的重要性。

原文摘要 · Abstract (English)

Conditional effects are commonly used measures for understanding how treatment effects vary across different groups, and are often used to target treatments/interventions to groups who benefit most. In this work we review existing methods and propose novel ones, focusing on the odds ratio (OR) and the risk ratio (RR). While estimation of the conditional average treatment effect (ATE) has been widely studied, estimators for the OR and RR lag behind, and cutting edge estimators such as those based on doubly robust transformations or orthogonal risk functions have not been generalized to these parameters. We propose such a generalization here, focusing on the DR-learner and the R-learner. We derive orthogonal risk functions for the OR and RR and show that the associated pseudo-outcomes satisfy second-order conditional-mean remainder properties analogous to the ATE case. We also evaluate estimators for the conditional ATE, OR, and RR in a comprehensive nonparametric Monte Carlo simulation study to compare them with common alternatives under hundreds of different data-generating distributions. Our numerical studies provide empirical guidance for choosing an estimator. For instance, they show that while parametric models are useful in very simple settings, the proposed nonparametric estimators significantly reduce bias and mean squared error in the more complex settings expected in the real world. We illustrate the methods in the analysis of physical activity and sleep trouble in U.S. adults using data from the National Health and Nutrition Examination Survey (NHANES). The results demonstrate that our estimators uncover substantial treatment effect heterogeneity that is obscured by traditional regression approaches and lead to improved treatment decision rules, highlighting the importance of data-adaptive methods for advancing precision health research.

因果推断治疗异质性正交学习精准医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。