梳理大模型智能体的安全保障体系,揭示当前技术的四大瓶颈。
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
- 构建三层次安全框架,整合规范、验证与执行
- 实测显示运行时监控可减少65%不安全操作
- 提出十大挑战,推动可信智能体发展
大模型智能体日益执行不可逆的真实世界操作,如数据库更新、API调用和文件操作。然而,现有系统无法为智能体生成的计划提供形式化支撑的任务级安全保证。研究分散在规范、验证与执行三个方向,限制了对现有方法优劣的理解。为此,我们基于PRISMA 2020标准,从六个学术数据库中检索2022至2026年发表的38项研究进行系统综述。分析揭示四个关键发现:第一,规范瓶颈仍是主要挑战,自然语言到形式化翻译的语义正确率仅24%至35%,削弱下游验证效果;第二,运行时监控是最成熟的执行策略,在受控环境中可减少40%至65%的不安全行为,但无法提供完全安全保证;第三,验证器税现象表明,阻止94%的不安全动作仍可能导致不足5%的任务成功完成,因智能体会绕道采用其他不安全路径;第四,尚无现有方法能同时满足可靠性、可扩展性、语义正确性和任务级安全保留。本文贡献包括三层次分类体系、现有技术的对比分析、验证器税的证据整合,以及一份十问题的研究议程,以推进可信自主智能体的发展。
原文摘要 · Abstract (English)
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。