arXiv:2607.03999stat.MEcs.LG2026-07

提出新方法,同时精准检测治疗差异并保证置信区间有效。

Significance-First Splitting: Aligning Treatment Heterogeneity Detection with Honest Estimation

  • 用显著性检验的平方t统计量做分裂,直接捕捉交互作用
  • 在三个合成数据上90%置信区间覆盖率接近名义水平
  • 适用于需要可靠推断的因果分析,如医疗或营销实验

估计异质性处理效应(CATE)需同时检测效应修饰和量化估计不确定性。现有基于树的方法存在权衡:基于显著性的方法(Radcliffe and Surry 2011)可直接识别亚组交互,但缺乏有效推断;诚实因果树(Athey and Imbens 2016)提供名义置信区间覆盖,但使用与结果无关的分裂标准,牺牲了交互敏感性。本文提出一种融合显著性分裂与诚实样本分割及交叉验证的混合算法。分裂准则采用治疗×侧边交互的平方t统计量(t²),当交互较强时,该准则与诚实的EMSE_τ准则直接对齐。事后诚实交叉验证选择代价-复杂度惩罚,生成单一原则性估计器,在叶节点层面实现名义置信区间覆盖。对于森林模型,保留自举计数向量以实现无穷小重抽样(IJ)方差估计,而非正式点推断。在Athey and Imbens(2016)的三个合成设计中,单棵树在90%名义水平下达到约90%的叶平均置信区间覆盖率(每设计200次重复);在Criteo、Hillstrom和Starbucks uplift数据集上,其Qini系数表现匹配S-learner、T-learner和GRF基线。本工作配套开源Python包,支持sklearn兼容接口、可复现种子和完整测试覆盖(https://codeberg.org/hadjipantelis/rattus)。

原文摘要 · Abstract (English)

Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty. Existing tree-based methods make an uneasy trade-off: significance-based approaches (Radcliffe and Surry 2011) identify subgroup interactions directly but lack valid inference; honest causal trees (Athey and Imbens 2016) deliver nominal confidence interval coverage but use outcome-agnostic splitting criteria that sacrifice interaction sensitivity. We introduce a hybrid algorithm that fuses significance-based splitting with honest sample-splitting and cross-validation. Our splitting criterion uses the squared $t$-statistic for the treatment $\times$ side interaction ($t^2$), which is shown to be directly aligned with the honest $\text{EMSE}_τ$ criterion when the interaction is strong. Post-hoc honest cross-validation selects the cost-complexity penalty, giving a single principled estimator with nominal CI coverage at the leaf level. For forests, we retain bootstrap count vectors to enable an infinitesimal jackknife (IJ) variance estimate of Monte-Carlo convergence rather than formal pointwise inference. On the three synthetic designs from (Athey and Imbens 2016) the single tree achieves approximately 90% leaf-average CI coverage at the 90% nominal level across all three designs (200 replications each); on the Criteo, Hillstrom and Starbucks uplift datasets we match Qini coefficient performance of S-, T-learner and GRF baselines. An open-source Python package with reproducible seeds, sklearn-compatible API, and full test coverage accompanies this work (https://codeberg.org/hadjipantelis/rattus).

因果推断异质性处理置信区间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。