arXiv:2502.15865q-fin.GNcs.AI2025-02被引 15

金融大模型代理需优先评估风险,而非仅看表现精度。

Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk

  • 从模型、流程到系统三层压力测试风险
  • 六款代理在高风险任务中暴露隐藏缺陷
  • 建议将安全预算作为核心成功指标

现有基准测试过度关注大语言模型(LLM)代理在金融领域的表现,却忽视其部署安全性。准确性指标和收益评分制造了虚假可靠性,忽略了幻觉事实、过时数据及对抗性提示攻击等漏洞。我们主张:金融LLM代理应首先评估其风险特征,而非点估计性能。基于风险工程原则,提出模型、工作流、系统三个层级的应力测试框架,用于模拟真实故障模式。通过审计六款基于API和开源权重的金融LLM代理在三项高影响力任务中的表现,揭示了传统基准未发现的潜在弱点。最后提出具体建议:未来研究应审计风险感知指标,发布测试场景与数据集,并将‘安全预算’列为首要成功标准。唯有重新定义‘好’的标准,才能推动金融领域负责任的AI发展。

原文摘要 · Abstract (English)

Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy. We argue that accuracy metrics and return-based scores provide an illusion of reliability, overlooking vulnerabilities such as hallucinated facts, stale data, and adversarial prompt manipulation. We take a firm position: financial LLM agents should be evaluated first and foremost on their risk profile, not on their point-estimate performance. Drawing on risk-engineering principles, we outline a three-level agenda: model, workflow, and system, for stress-testing LLM agents under realistic failure modes. To illustrate why this shift is urgent, we audit six API-based and open-weights LLM agents on three high-impact tasks and uncover hidden weaknesses that conventional benchmarks miss. We conclude with actionable recommendations for researchers, practitioners, and regulators: audit risk-aware metrics in future studies, publish stress scenarios alongside datasets, and treat ``safety budget'' as a primary success criterion. Only by redefining what ``good'' looks like can the community responsibly advance AI-driven finance.

大模型代理金融AI风险评估安全测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。