用双变量贝塔分布建模干预效果,推断客户是否真会流失
Identifying counterfactual probabilities using bivariate distributions and uplift modeling
- 基于提升模型预测得分,拟合双变量贝塔分布
- 可估计客户在有无营销时的联合流失概率
- 无需额外因果假设,适合电信客户留存分析
提升模型通过治疗组与对照组潜在结果的差异来估计干预的因果效应,而反事实识别旨在恢复这些潜在结果的联合分布(例如:‘如果未提供营销优惠,这位客户还会流失吗?’)。该联合反事实分布比单一提升值包含更多信息,但更难估计。然而,两者具有协同性:可利用提升模型辅助反事实估计。本文提出一种反事实估计器,将预测的提升分数拟合为双变量贝塔分布,从而得到反事实结果的后验分布。该方法仅依赖于提升模型的因果假设,无需额外前提。模拟实验验证了其有效性,可用于电信行业客户流失问题,揭示标准机器学习或提升模型无法获取的深层洞察。
原文摘要 · Abstract (English)
Uplift modeling estimates the causal effect of an intervention as the difference between potential outcomes under treatment and control, whereas counterfactual identification aims to recover the joint distribution of these potential outcomes (e.g., "Would this customer still have churned had we given them a marketing offer?"). This joint counterfactual distribution provides richer information than the uplift but is harder to estimate. However, the two approaches are synergistic: uplift models can be leveraged for counterfactual estimation. We propose a counterfactual estimator that fits a bivariate beta distribution to predicted uplift scores, yielding posterior distributions over counterfactual outcomes. Our approach requires no causal assumptions beyond those of uplift modeling. Simulations show the efficacy of the approach, which can be applied, for example, to the problem of customer churn in telecom, where it reveals insights unavailable to standard ML or uplift models alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。