用少样本实现接近最优的治疗分配,关键在于粗略估计即可。
Good Allocations from Bad Estimates
- 用粗略估计替代精确估计,实现近优分配
- 仅需 $O(M/ε)$ 样本,比传统方法快一倍
- 适合样本稀缺的真实随机试验场景
条件平均处理效应(CATE)估计是针对异质人群进行治疗靶向的黄金标准。该方法在 $M$ 个不同人群分层中,将处理效应估计误差控制在 $ε>0$ 以内,按估计值降序分配治疗直至预算耗尽。通常此方法需要 $O(M/ε^2)$ 个样本。若目标为所有处理效应均达到 $ε$ 精度,这是最优的。本文表明,在处理效应分布自然的情况下,仅需 $O(M/ε)$ 样本即可实现与 CATE 相同的总治疗效果。核心洞察是:粗略估计已足够实现近优分配。此外,预算灵活性可进一步降低样本复杂度。我们在多个真实世界随机对照试验(RCT)数据集上评估了该算法,在所有情况下均以极少样本找到近乎最优的治疗分配。本工作揭示了处理效应估计与分配之间的根本差异:后者所需样本远少于前者。
原文摘要 · Abstract (English)
Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects up to an error $ε> 0$ in each of $M$ different strata of the population, targeting individuals in decreasing order of estimated treatment effect until the budget runs out. In general, this method requires $O(M/ε^2)$ samples. This is best possible if the goal is to estimate all treatment effects up to an $ε$ error. In this work, we show how to achieve the same total treatment effect as CATE with only $O(M/ε)$ samples for natural distributions of treatment effects. The key insight is that coarse estimates suffice for near-optimal treatment allocations. In addition, we show that budget flexibility can further reduce the sample complexity of allocation. Finally, we evaluate our algorithm on various real-world RCT datasets. In all cases, it finds nearly optimal treatment allocations with surprisingly few samples. Our work highlights the fundamental distinction between treatment effect estimation and treatment allocation: the latter requires far fewer samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。