提出新框架,让反欺诈系统自动判断何时可安全授权,避免用过时数据误判。
When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload

- 基于证据新鲜度与人力容量,动态决定哪些欺诈案件可自动处理
- 在三个真实数据集上实现超80%自动化率,同时将人工审查量降至46%以下
- 揭示审计频率与工作量的权衡关系,适合风控系统设计者参考
欺诈操作需在自动批准、人工复核和自动拦截间分配任务,但评估这些操作所需的标签具有选择性且延迟。预测得分虽能排序案件,却无法判断证据是否最新且具有代表性。本文提出新鲜度约束审计能力(FCAC)框架,将自动化视为受行动风险、证据新鲜度及共享复核容量限制的授权决策。该框架通过成熟随机审计结果与预设时间窗口评估候选行动区域:支持区域实现自动化,不支持区域保留人工复核。最终决策记录包含证据年龄、审计需求、总复核工作量、风险暴露及兼容的时间变化。研究表明,在未限制未观测标签演变的情况下,当前行动风险无法识别。在代表性随机审计、标签无关证据窗口及历史与当前风险关联条件下,可实现对不安全授权的联合有限样本控制。在IEEE-CIS、ULB-Worldline和Elliptic++上的模拟审计显示,零漂移自动化率分别为84.4%、67.4%、81.3%,总复核工作量为24.1%、46.0%、43.1%。实验揭示审计容量权衡:稀疏审计延后授权,密集审计最终增加工作量。独立的BAF压力测试进一步表明,回退阈值应基于候选证据特性,而非统一的风险上限比例。研究指出,审计新鲜度与分析师容量是欺诈决策支持系统的共同设计要素。
原文摘要 · Abstract (English)
Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these actions are selective and delayed. Predictive scores order cases, but they do not show whether the evidence is current and representative enough to delegate an action to the model. We develop freshness-constrained audit capacity (FCAC), a decision-support framework that treats automation as an authorization decision constrained by action risk, evidence freshness, and shared review capacity. It evaluates candidate action regions from mature randomized audits and a prespecified temporal allowance. Supported regions are automated; unsupported regions remain in review. The resulting decision record reports evidence age, audit demand, total review workload, value exposure, and compatible temporal change. We show that current action risk is unidentified without restricting unobserved label evolution. Under representative randomized audits, label-independent evidence windows, and a prespecified condition linking historical and current action risk, we derive simultaneous finite-sample control of unsafe authorization. Chronological evaluations with simulated audits on IEEE-CIS, ULB-Worldline, and Elliptic++ yield zero-drift automation rates of 84.4%, 67.4%, and 81.3%, with total review workloads of 24.1%, 46.0%, and 43.1%. The experiments reveal an audit-capacity trade-off: sparse auditing delays authorization, whereas intensive auditing eventually increases workload. A separately specified BAF stress test further indicates that fallback thresholds must reflect candidate-specific evidence rather than a common fraction of the risk limit. These findings identify audit freshness and analyst capacity as joint design considerations for fraud decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。