测试银行智能体在复杂欺诈攻击下的安全防护能力。
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

- 构建可执行的对抗性测试框架,模拟真实对话中的身份欺骗与权限滥用。
- 107个公开攻击场景中,四款智能体防御成功率仅49%至65%。
- 适合研究金融AI安全、模型鲁棒性或反欺诈系统的团队使用。
对话代理如今可通过工具访问客户数据库和内部政策文档,直接处理账户信息变更、密码重置或资金转移等操作,使客户服务与授权、防欺诈及合规审查紧密绑定。现有金融欺诈基准多针对静态交易或消息分类,通用安全基准则聚焦提示注入或通用有害行为,均未评估政策驱动型银行代理在用户通过对话逐步操纵身份、权限与信任时的安全表现。本文提出FraudBench,基于τ²-bench双控框架与τ-Knowledge银行环境构建可执行基准。代理与模拟攻击者通过共享可变账户状态和工具交互,代理可授予特定工具访问权;环境包含698份内部政策文档供代理检索。FraudBench包含150个精心设计的对抗性场景,其中107个公开基准场景(涵盖90种欺诈机制及17个链式自适应攻击)用于统一评测,另有43个链式攻击作为保留测试集。安全性具有历史依赖性:单控任务满足所有前提条件但一项例外,而自适应攻击通过前期探测、承认或失败尝试,使后续看似合理的请求变得不安全。每个场景标注可观测证据、禁止行为、安全处置方式及干预点。对四款代理的初步单次评估显示,在107个评分任务中攻击成功率介于49%至65%,其中“资金洗钱”与“首方欺诈”为跨模型最常见的弱点。
原文摘要 · Abstract (English)
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $τ^2$-bench dual-control framework and the $τ$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。