arXiv:2605.27474stat.MLcs.LG2026-05

针对极端事件,提出可输出尾部特征的因果估计新方法

Stop Suppressing the Tail: Causal Inference for Extreme Events

论文配图:Stop Suppressing the Tail: Causal Inference for Extreme Events
图 1 · 摘自论文原文
  • 用中心化中位数分解残差,打破尾部推断的循环依赖
  • 在重尾数据上,极端事件预测误差降低11%~25.5%
  • 适合金融、气候等高风险领域,支持极端值外推拒绝机制

估计连续处理对结果的影响(平均剂量响应函数,ADRF)是因果推断的核心任务。然而,在结果具有重尾分布时,标准鲁棒双机器学习(DML)会刻意抑制极端值以稳定总体均值。在金融收益或气候损失等高风险场景中,这些1/1000的极端事件正是关键目标。现有从模型残差读取尾部的方法存在循环依赖,导致尾部形状推断随核心估计器(如Huber与Welsch)切换而剧烈波动。本文提出一种ADRF估计器,可输出结构化的尾部形状结果。其尾部诊断(PDHTE+JK)基于中心化中位数后的残差,成功打破循环依赖,使诊断结果不受核心方法选择影响。输出包含四类治疗条件量:尾部形状ξ̂(t)、深尾部重现水平Q̂α(t)、条件短缺Ŝα(t)、恢复后的均值ADRF,以及在数据不支持极端值建模时的显式外推拒绝机制。相比核加权分位数回归(QR),该方法在重尾面板数据上将深尾部(α=0.001)重现水平的平均绝对误差(MAE)降低11%,条件短缺的MAE降低25.5%;在样本稀缺情形(n≤2000)下,MAE降低20%-29%。在freMTPL2车险索赔数据上,该方法成功在对数索赔尺度触发显式外推拒绝,而QR和仅损失型DML无法实现。

原文摘要 · Abstract (English)

Estimating how an outcome responds to a continuous treatment (the Average Dose-Response Function, or ADRF) is a core causal-inference primitive. However, when outcomes possess heavy tails, standard robust double machine learning (DML) deliberately suppresses these extremes to stabilize the bulk average. In high-stakes settings, such as financial returns or climate losses, this omitted 1-in-1000 extreme event is the actual target quantity. Furthermore, current methods that read the tail from a model's residuals suffer from circular dependence, causing tail shape inferences to shift drastically based solely on whether the core estimator is switched between Huber and Welsch. The research proposes an ADRF estimator that emits a structured tail-shape output alongside the standard point estimate. Its tail diagnostic (PDHTE+JK) evaluates the per-treatment tail shape from the outcome centered by a pilot median, successfully breaking the circular dependence and rendering the diagnostic invariant to the choice of core method. The output encompasses four treatment-conditional quantities: tail shape $\hatξ(t)$, deep-tail return levels $\hat{Q}_α(t)$, conditional shortfalls $\hat{S}_α(t)$, the recovered mean ADRF, and an explicit refusal mechanism that declines extrapolation when extreme-value modeling is unsupported by the data. Compared to kernel-weighted quantile regression (QR), the proposed estimator reduces deep-tail ($α=0.001$) return-level MAE by 11% and conditional-shortfall MAE by 25.5% across a heavy-tailed panel. It also achieves a 20-29% MAE reduction in sample-scarce regimes ($n\le2000$). On freMTPL2 motor-insurance claims, it successfully triggered an explicit extrapolation refusal on the log-claim scale, which neither QR nor loss-only DML can produce.

因果推断重尾分析极端事件外推拒绝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。