arXiv:2601.07853cs.CRcs.AI2026-01被引 6

首个面向金融智能体的实战安全评测基准,揭示现有防护漏洞。

FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments

  • 构建31个真实金融场景的可操作测试环境,模拟状态变更与合规约束。
  • 实测顶尖模型攻击成功率高达50%,最稳健系统仍达6.7%。
  • 专为金融领域设计,适合研究者和开发者评估智能体安全性。

由大语言模型驱动的金融代理在投资分析、风险评估和自动化决策中日益普及,其规划、调用工具及修改可变状态的能力带来了高风险、高监管环境下的新安全挑战。然而,现有安全评估多聚焦于语言模型内容合规或抽象代理设置,无法捕捉实际操作流程中因状态变更引发的安全风险。为此,我们提出FinVault,首个面向金融代理的执行基础安全基准,包含31个基于监管案例的沙盒场景,配备可写入状态的数据库与明确合规约束,并涵盖107个真实世界漏洞及963个测试用例,系统覆盖提示注入、越狱攻击、金融适配攻击以及良性输入以评估误报率。实验表明,现有防御机制在真实金融代理环境中仍无效,最先进模型平均攻击成功率达50.0%,即使最鲁棒系统仍有6.7%的攻击成功率,凸显当前安全设计迁移能力有限,亟需更强的金融专用防护方案。代码开源地址:https://github.com/aifinlab/FinVault。

原文摘要 · Abstract (English)

Financial agents powered by large language models (LLMs) are increasingly deployed for investment analysis, risk assessment, and automated decision-making, where their abilities to plan, invoke tools, and manipulate mutable state introduce new security risks in high-stakes and highly regulated financial environments. However, existing safety evaluations largely focus on language-model-level content compliance or abstract agent settings, failing to capture execution-grounded risks arising from real operational workflows and state-changing actions. To bridge this gap, we propose FinVault, the first execution-grounded security benchmark for financial agents, comprising 31 regulatory case-driven sandbox scenarios with state-writable databases and explicit compliance constraints, together with 107 real-world vulnerabilities and 963 test cases that systematically cover prompt injection, jailbreaking, financially adapted attacks, as well as benign inputs for false-positive evaluation. Experimental results reveal that existing defense mechanisms remain ineffective in realistic financial agent settings, with average attack success rates (ASR) still reaching up to 50.0\% on state-of-the-art models and remaining non-negligible even for the most robust systems (ASR 6.7\%), highlighting the limited transferability of current safety designs and the need for stronger financial-specific defenses. Our code can be found at https://github.com/aifinlab/FinVault.

金融智能体安全评测大模型安全提示攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。