融合图特征与智能调查,提升金融欺诈检测的可审计性。
Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation

- 构建多层检测管道,整合图结构特征与模型解释。
- 图特征在中间评分样本中显著提升欺诈识别率,漏检率降低25%。
- 调查代理虽能生成合理解释,但决策准确率低于直接阈值法。
欺诈检测系统需应对交易量增长,同时保持可解释性和可审查性。本研究在PaySim数据集上构建多层流水线,结合梯度提升分类器、图结构特征、基于自编码器的异常信号、TreeSHAP解释,以及一个受控的LLM调查代理,用于处理分类器输出不确定的案例。在修正模拟器特有的平衡偏差后,图特征与异常信号未提升全集平均精度,但在中间评分样本中显著改善欺诈排序。在注入多账户欺诈环的实验中,工程化图特征捕获全部测试交易,而基线方法遗漏约四分之一。调查代理在60个样本上准确率为65.0%,低于分类器阈值法的71.7%。其8次决策更改中,6次将正确判断误判为错误,且每次均生成合理理由。基于分歧的升级规则成功标记了两次错误,未误标任何正确判断。结论:各组件仅在特定条件下有效,代理的合理解释不等于更优决策。
原文摘要 · Abstract (English)
Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, an autoencoder-based anomaly signal, TreeSHAP explanations, and a bounded LLM investigation agent applied to cases the classifier scores uncertainly. Before any model comparison, we identify and remove a simulator-specific balance shortcut that would otherwise inflate baseline performance. After this correction, neither the graph features nor the anomaly signal improves Average Precision on the full test set. Both, however, rank fraud better within the subset of cases receiving intermediate baseline scores. In a controlled experiment with injected multi-account fraud rings, engineered structural features recover all injected test transactions, while the tabular baseline misses roughly a quarter of them. The investigation agent underperforms direct thresholding of the classifier it relies on, reaching 65.0% accuracy against 71.7% on a balanced 60-case sample, despite having access to model explanations, graph context, and retrieved reference cases. Of the eight decisions the agent changed, six replaced correct classifier outputs with errors, and it produced a coherent written rationale in each case. An exploratory disagreement-based escalation rule flagged two of these agent errors for human review without flagging any correct decision. We conclude that each component of a layered fraud system contributes only under specific conditions, and that a plausible rationale from an investigation agent is not evidence of a better decision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。