arXiv:2608.24551cs.LGcs.AI2026-08

金融风控模型抗攻击能力评估需考虑协议差异,否则结论可能失真。

FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment

论文配图:FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment
图 1 · 摘自论文原文
  • 设计三种不同约束条件的评测协议,对比评估结果差异
  • 同一数据下,攻击方式不同导致可攻击样本数相差近千倍
  • 适合关注金融模型安全性的研究者与开发者参考

机器学习模型广泛用于金融欺诈与信用风险检测,但其对抗鲁棒性难以评估,因金融表格数据存在领域特定约束、严重类别不平衡及不对称攻击能力。本文提出 FraudBench,一个协议敏感的金融风控对抗鲁棒性基准。不同于将领域约束作为事后验证,FraudBench 在三个匹配协议下评估相同模型-攻击-防御设置:无约束攻击、事后可行性过滤、部署感知的约束融合攻击。覆盖四个公开金融数据集,评估神经网络、树模型与集成模型在三种攻击设定下的表现。结果显示,鲁棒性结论高度依赖评测协议。在 Lending Club Loan Data 白盒设置下,事后过滤平均仅保留 3.7 个可行扰动样本,而攻击中投影结合攻击者可变性掩码则产生 2,832.3 个可行扰动样本,扰动预算相同。IEEE-CIS 结果表明可行性与攻击能力是独立维度,黑盒评估显示协议选择可改变模型族排名。结论建议:评估应同时报告预测性能下降与攻击可行性,并将领域约束融入攻击生成过程,而非事后处理。

原文摘要 · Abstract (English)

Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. We argue that, in this setting, robustness is not only an attribute of the model, but also an attribute of the evaluation protocol. Different ways of enforcing constraints and capability can lead to substantially different robustness conclusions. This paper presents FraudBench, a protocol-sensitive benchmark for adversarial robustness evaluation in financial fraud and credit-risk detection. Rather than treating domain constraints as post-hoc validity checks, FraudBench evaluates the same dataset--model--attack--defence setting under three matched protocols: unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks. FraudBench covers four public financial datasets, and evaluates neural, tree-based, and ensemble models using three attack settings. Our results show that robustness conclusions are highly protocol-sensitive. On Lending Club Loan Data under the white-box setting, post-hoc filtering leaves only 3.7 feasible-flipped examples on average, whereas in-attack projection with attacker mutability masking produces 2,832.3 feasible-flipped examples under the same perturbation budget. The results on IEEE-CIS further show that feasibility and attacker capability are separate axes, while black-box evaluation shows that protocol choice can alter model-family rankings. These findings suggest that fraud robustness evaluation should report predictive degradation and attack feasibility jointly, and should incorporate domain constraints into attack generation rather than treating them as post-processing checks.

金融风控对抗鲁棒性评测基准数据约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。