fraud检测需按不同观察机制分类,否则效率严重下降
Fraud Type Decomposition and the Observation-Mechanism Taxonomy:Class-Specific Detection Limits in Payment Networks
- 将欺诈分为五类,每类对应不同标记流程
- 分类估计比合并估计更高效,差距由詹森不等式决定
- 揭示各类别检测上限,适合反欺诈系统设计者
支付网络中的欺诈检测依赖于异构且不完美的观测过程生成的标签,但现有方法将欺诈视为同质的二元变量。我们证明这一假设在结构上错误,并导致可证明的低效。本文提出一种观察机制分类法,将欺诈划分为五类,每类具有独特的审查与标注流程。我们证明:按类别分别估计欺诈率并聚合,严格优于合并估计,其效率差距由异质观测率引发的詹森惩罚决定。针对每类,我们推导出检测的理论制约因素,包括内生标签污染、结构性不可观测性及特征无效性。结果表明,欺诈检测本质上是多个由各自观测结构和检测极限支配的独立估计问题。
原文摘要 · Abstract (English)
Fraud detection in payment networks relies on labels generated through heterogeneous and imperfect observation processes, yet existing approaches treat fraud as a homogeneous binary variable. We show that this assumption is structurally incorrect and leads to provable inefficiency. We introduce an observation-mechanism taxonomy that partitions fraud into five classes, each defined by a distinct censorship and labeling pipeline. We prove that estimating fraud rates separately by class and aggregating strictly dominates pooled estimation, with the efficiency gap characterized as a Jensen penalty arising from heterogeneous observation rates. For each class, we derive the binding theoretical constraint on detection, including endogenous label corruption, structural non-observability, and feature non-informativeness. These results establish that fraud detection is fundamentally a collection of distinct estimation problems, each governed by its own observation structure and detection limit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。