首个系统评估大模型代理安全性的基准,发现16个主流代理安全得分均低于60%。
Agent-SafetyBench: Evaluating the Safety of LLM Agents
- 构建包含2000个测试用例的综合性安全评测框架
- 16个主流代理在8类风险上平均安全得分不足60%
- 揭示代理缺乏鲁棒性与风险意识,提示需更深层防御策略
随着大语言模型(LLMs)越来越多地作为智能体部署,其在交互环境和工具使用中的集成带来了超越模型本身的新安全挑战。然而,缺乏全面的评测基准严重阻碍了对代理安全性的有效评估与改进。本文提出Agent-SafetyBench,一个用于评估LLM代理安全性的综合性基准。该基准涵盖349个交互环境与2000个测试用例,评估8类安全风险,并覆盖10种常见不安全交互中的失效模式。对16个主流LLM代理的评估显示:无一代理的安全得分超过60%,凸显当前代理存在显著安全缺陷。通过失效模式与有用性分析,我们总结出两类根本性安全问题:缺乏鲁棒性与缺乏风险意识。此外,研究发现仅依赖防御提示可能不足以解决这些安全问题,强调需要更先进、更稳健的应对策略。为推动该领域进展,Agent-SafetyBench已开源至https://github.com/thu-coai/Agent-SafetyBench/,以促进代理安全性评测与提升的进一步研究。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed as agents, their integration into interactive environments and tool use introduce new safety challenges beyond those associated with the models themselves. However, the absence of comprehensive benchmarks for evaluating agent safety presents a significant barrier to effective assessment and further improvement. In this paper, we introduce Agent-SafetyBench, a comprehensive benchmark designed to evaluate the safety of LLM agents. Agent-SafetyBench encompasses 349 interaction environments and 2,000 test cases, evaluating 8 categories of safety risks and covering 10 common failure modes frequently encountered in unsafe interactions. Our evaluation of 16 popular LLM agents reveals a concerning result: none of the agents achieves a safety score above 60%. This highlights significant safety challenges in LLM agents and underscores the considerable need for improvement. Through failure mode and helpfulness analysis, we summarize two fundamental safety defects in current LLM agents: lack of robustness and lack of risk awareness. Furthermore, our findings suggest that reliance on defense prompts alone may be insufficient to address these safety issues, emphasizing the need for more advanced and robust strategies. To drive progress in this area, Agent-SafetyBench has been released at https://github.com/thu-coai/Agent-SafetyBench/ to facilitate further research in agent safety evaluation and improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。