解决支付网络欺诈标签偏差问题,实现最优检测性能。
Causal Label Recovery in Payment Networks
- 提出顺序三重稳健估计器,同时校正四类标签缺陷
- 理论证明达到半参数效率下界,比传统方法更优
- 可使用数天旧数据训练,显著提升模型时效性
支付网络中的欺诈检测模型依赖于存在系统性偏差的退款标签。每个标签需通过三个连续环节:授权(被拒交易不生成标签)、发卡行报告(未报告的欺诈无法被发现)以及延迟(待处理退款在训练时缺失)。幸存的标签还可能因第一方滥用或发卡行误判而受损。前序论文[arXiv:2605.27557]证明这四类缺陷导致检测性能的极小极大下界。本文提出问题:能否达到该下界?我们将观测流程建模为具有三个倾向性阶段和一个污染层的序列缺失数据问题,构造出顺序三重稳健(STR)估计器。该估计器可同时纠正四类缺陷,并达到半参数效率界——无其他估计器能拥有更低渐近方差。其具备顺序三重稳健性:每一步只需倾向性模型或结果回归之一正确,无需两者均正确。通过噪声率调整伪标签实现污染校正,采用经验贝叶斯收缩稳定小发卡行的逆倾向权重,引入插值方差估计量获得有效置信区间,并基于伯恩斯坦集中不等式提供有限样本保证。运营层面,推导出最优训练延迟——最小化标签质量损失与模型陈旧性的成熟窗口,证明STR允许使用数天而非数月的老数据训练,使模型新鲜度脱离退款成熟周期。无论样本量大小,STR在均方误差上严格优于朴素退款训练方法。
原文摘要 · Abstract (English)
Fraud detection models in payment networks train on chargeback labels that are systematically biased. Every label must survive three sequential gates: authorization (declined transactions generate no labels), issuer reporting (unreported fraud is invisible), and delay (pending chargebacks are missing at training time). Labels that do arrive may be corrupted by first-party misuse or issuer misclassification. A companion paper [arXiv:2605.27557] proved that these four impairments impose a minimax lower bound on detection performance. This paper asks: can that bound be achieved? We formalize the observation pipeline as a sequential missing-data problem with three propensity stages and a corruption layer, and construct the Sequential Triply Robust (STR) estimator. The STR corrects for all four impairments simultaneously and achieves the semiparametric efficiency bound -- no estimator can have lower asymptotic variance. It is sequentially triply robust: at each gate, consistency requires only that either the propensity model or the outcome regression is correctly specified, not both. We provide corruption correction via noise-rate-adjusted pseudo-labels, empirical Bayes shrinkage to stabilize inverse-propensity weights for small issuers, a plug-in variance estimator yielding valid confidence intervals, and a Bernstein concentration inequality for finite-sample guarantees. On the operational side, we derive the optimal training delay -- the maturity window that minimizes the sum of label-quality loss and model staleness -- and prove that the STR permits training on data that is days old rather than months old, decoupling model freshness from the chargeback maturity cycle. The STR provably dominates naive chargeback-based training in mean squared error for any sample size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。