arXiv:2602.20383stat.MEcs.LG2026-02

机器学习预测个体治疗效果时,群体聚合会引入系统性偏差,本文提出检测与修正方法。

Detecting and Mitigating Group Bias in Heterogeneous Treatment Effects

  • 通过比较模型推断与实验观测的群体平均效应,定义并检测群体偏差
  • 提出可闭式求解的收缩校正方法,有效降低偏差,提升决策准确性
  • 适用于个性化投放、医疗干预等需群体公平性的场景

在随机实验中,即使个体条件平均处理效应(CATE)模型正确设定且训练充分,将预测的CATE聚合到群体层面也无法一般性地恢复群体平均处理效应(GATE),导致系统性偏差。本文建立统一统计框架,定义群体偏差为模型推断的GATE与实验识别的GATE之差,导出渐近正态估计量,并提供简便的统计检验方法。针对缓解,提出基于收缩的偏差校正方法,理论最优解与实际可行解均具闭式表达。该框架假设最少,仅需计算样本矩,适用于各类随机实验数据。通过分析偏差修正对利润最大化个性化投放的经济影响,揭示何时改变投放策略及利润权衡。在大型数字平台的实验数据上验证了理论结果与实证表现。

原文摘要 · Abstract (English)

Heterogeneous treatment effects (HTEs) are increasingly estimated using machine learning models that produce highly personalized predictions of treatment effects. In practice, however, predicted treatment effects are rarely interpreted, reported, or audited at the individual level but, instead, are often aggregated to broader subgroups, such as demographic segments, risk strata, or markets. We show that such aggregation can induce systematic bias of the group-level causal effect: even when models for predicting the individual-level conditional average treatment effect (CATE) are correctly specified and trained on data from randomized experiments, aggregating the predicted CATEs up to the group level does not, in general, recover the corresponding group average treatment effect (GATE). We develop a unified statistical framework to detect and mitigate this form of group bias in randomized experiments. We first define group bias as the discrepancy between the model-implied and experimentally identified GATEs, derive an asymptotically normal estimator, and then provide a simple-to-implement statistical test. For mitigation, we propose a shrinkage-based bias-correction, and show that the theoretically optimal and empirically feasible solutions have closed-form expressions. The framework is fully general, imposes minimal assumptions, and only requires computing sample moments. We analyze the economic implications of mitigating detected group bias for profit-maximizing personalized targeting, thereby characterizing when bias correction alters targeting decisions and profits, and the trade-offs involved. Applications to large-scale experimental data at major digital platforms validate our theoretical results and demonstrate empirical performance.

因果推断群体偏差个性化推荐机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。