提出长效安全防护框架,动态识别并抵御大模型代理的各类风险。
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
- 通过自适应生成安全检测机制,实时响应任务变化
- 在多个任务中显著降低系统性与任务特定风险
- 兼容多种大模型代理,适合长期部署的安全需求
大语言模型(LLMs)在动态环境中被广泛用作自主代理以完成复杂任务,展现出强大的问题解决能力和场景适应性。然而,其作为代理使用也带来了显著风险:任务特定风险由代理管理员根据具体任务要求和约束定义;系统性风险则源于设计缺陷或交互漏洞,可能破坏信息的机密性、完整性或可用性(CIA),引发安全威胁。现有防御机制难以有效且自适应地应对这些风险。本文提出AGrail,一种长效代理安全防护框架,具备自适应安全检测生成、高效优化及工具兼容灵活等特性。大量实验表明,AGrail不仅在应对任务特定风险和系统性风险方面表现优异,还具备跨不同大模型代理任务的可迁移能力。
原文摘要 · Abstract (English)
The rapid advancements in Large Language Models (LLMs) have enabled their deployment as autonomous agents for handling complex tasks in dynamic environments. These LLMs demonstrate strong problem-solving capabilities and adaptability to multifaceted scenarios. However, their use as agents also introduces significant risks, including task-specific risks, which are identified by the agent administrator based on the specific task requirements and constraints, and systemic risks, which stem from vulnerabilities in their design or interactions, potentially compromising confidentiality, integrity, or availability (CIA) of information and triggering security risks. Existing defense agencies fail to adaptively and effectively mitigate these risks. In this paper, we propose AGrail, a lifelong agent guardrail to enhance LLM agent safety, which features adaptive safety check generation, effective safety check optimization, and tool compatibility and flexibility. Extensive experiments demonstrate that AGrail not only achieves strong performance against task-specific and system risks but also exhibits transferability across different LLM agents' tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。