提出抗重尾异常值的双机器学习方法,显著提升剂量反应函数估计稳定性。
SHIFT: Robust Double Machine Learning for Average Dose-Response Functions under Heavy-Tailed Contamination

- 用加权Welsch损失结合渐进非凸优化,降低异常值影响
- 在25%局部污染下均方误差从1.03降至0.33,保持清洁数据性能
- 可识别异常样本,适合高风险因果推断场景
基于核加权局部线性平滑的平均剂量反应函数双机器学习方法易受异常值影响。本文提出SHIFT(自校准重尾内点拟合与调温法),结合交叉拟合、正交化与梯度非凸优化的核局部Welsch损失第二阶段,并采用基于后GNC残差中位绝对偏差的内点阈值进行防御性OLS重拟合。在局部污染率p=0.25的压力测试中,该设计使水平-均方误差由1.03降至0.33,而干净及均匀污染情况不变。在1400次主循环中,SHIFT在最坏情况下形状恢复的均方误差为0.325(仅次于Huber-DML的0.276);三者中最坏误差低于0.35时,仅SHIFT产生非均匀样本权重,对高斯跳跃生成过程的真异常值掩码平均F1达0.96(范围0.945–0.968)。配套六种极值理论诊断工具(Hill、GPD-MLE/PWM、GEV、均超额、参数稳定性、因果尾部系数),帮助区分Frechet与Weibull尾部类型,指导选择SHIFT或L1替代方案。扩展至二元处理CATE(Huber伪结果X-Learner)和时间序列ADRF(块交叉验证+滚动中位绝对偏差)。反直觉消融实验显示:在均匀污染下,线性扰动模型(Ridge、Lasso)优于梯度提升模型,颠覆了更灵活更好的常规认知。
原文摘要 · Abstract (English)
Double-machine-learning pipelines for the Average Dose-Response Function rely on kernel-weighted local-linear smoothers, which inherit unbounded functional influence: a single outlier within a kernel window biases the curve across the entire window. We introduce SHIFT (Self-calibrated Heavy-tail Inlier-Fit with Tempering), a robust DML estimator combining cross-fit nuisance orthogonalization with a kernel-local Welsch-loss second stage optimized by Graduated Non-Convexity, and -- the principal design choice -- a defensive OLS refit whose inlier cutoff is scaled by post-GNC residual MAD rather than the raw-outcome MAD. On a localized-contamination stress test at $p=0.25$ this design choice drops level-RMSE from 1.03 to 0.33 while leaving clean and uniformly-contaminated runs unchanged. Across 1,400 main-sweep fits, SHIFT has competitive worst-case shape recovery (RMSE $0.325$ at $p=0.25$, second to Huber-DML's $0.276$); among the three methods with worst-case RMSE below $0.35$, only SHIFT emits a non-uniform per-sample weight vector, recovering the ground-truth outlier mask at mean $F_1 \approx 0.96$ (range $0.945$--$0.968$) on Gaussian-jump DGPs. We pair the estimator with a six-technique Extreme Value Theory diagnostic suite (Hill, GPD-MLE/PWM, GEV, Mean Excess, parameter stability, causal tail coefficient) that lets a practitioner distinguish Frechet from Weibull regimes and choose between SHIFT and L1 alternatives on empirical grounds. Extensions to binary-treatment CATE (Huber pseudo-outcome X-Learner) and time-series ADRF (block-CV + rolling MAD) are included. A counter-intuitive ablation: linear nuisance models (Ridge, Lasso) outperform gradient-boosted nuisances for robust DML under uniform contamination, inverting the usual more-flexible-is-better heuristic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。