arXiv:2601.15360stat.MLcs.LG2026-01

针对数据不平衡与长尾分布,提出抗异常值的X-Learner模型。

Robust X-Learner: Breaking the Curse of Imbalance and Heavy Tails via Robust Cross-Imputation

  • 用γ-散度替代均方误差,抑制极端值干扰
  • 在Criteo数据集上使估计误差降低98.6%
  • 适合广告、医疗等存在极端样本的工业场景

在广告技术和医疗等工业应用中,估计异质处理效应(HTE)面临双重挑战:极端类别不平衡和重尾结果分布。尽管X-Learner框架通过交叉插补有效缓解不平衡问题,但其依赖均方误差(MSE)最小化时,会因少数极端观测值(“鲸鱼”)导致“异常值扩散”——将偏误传播至多数群体,破坏处理效应结构。为此,我们提出鲁棒X-Learner(RX-Learner),将红降γ-散度目标函数(在高斯假设下等价于Welsch损失)融入梯度提升机制,并基于极大化-极小化(MM)原则引入代理海森矩阵稳定非凸优化。在半合成Criteo Uplift数据集上的实证评估表明,与标准X-Learner相比,RX-Learner将异质效应估计精度(PEHE)降低了98.6%,成功将稳定“核心”群体与波动“边缘”群体解耦。

原文摘要 · Abstract (English)

Estimating Heterogeneous Treatment Effects (HTE) in industrial applications such as AdTech and healthcare presents a dual challenge: extreme class imbalance and heavy-tailed outcome distributions. While the X-Learner framework effectively addresses imbalance through cross-imputation, we demonstrate that it is fundamentally vulnerable to "Outlier Smearing" when reliant on Mean Squared Error (MSE) minimization. In this failure mode, the bias from a few extreme observations ("whales") in the minority group is propagated to the entire majority group during the imputation step, corrupting the estimated treatment effect structure. To resolve this, we propose the Robust X-Learner (RX-Learner). This framework integrates a redescending γ-divergence objective -- structurally equivalent to the Welsch loss under Gaussian assumptions -- into the gradient boosting machinery. We further stabilize the non-convex optimization using a Proxy Hessian strategy grounded in Majorization-Minimization (MM) principles. Empirical evaluation on a semi-synthetic Criteo Uplift dataset demonstrates that the RX-Learner reduces the Precision in Estimation of Heterogeneous Effect (PEHE) metric by 98.6% compared to the standard X-Learner, effectively decoupling the stable "Core" population from the volatile "Periphery".

因果推断异常值鲁棒重尾分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。