用智能体自身能力自动检测并约束危险工具调用流程。
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
- 利用大模型编排器的工具理解与执行能力,自动生成真实工作流。
- 在实际执行中验证出潜在危险流程,并生成安全约束规则。
- 适合关注大模型智能体安全性的研究人员与开发者使用。
将工具调用集成到大语言模型(LLMs)中,使智能体具备现实世界影响能力。然而,与独立的LLM不同,被攻陷的智能体可利用工具调用能力执行恶意工作流,造成更严重后果。我们提出AgentGuard框架,可自主发现并验证不安全的工具使用流程,随后生成安全约束以限制智能体行为,实现部署时的安全基线保障。AgentGuard利用LLM编排器固有的能力——对工具功能的理解、可扩展且真实的流程生成能力、以及工具执行权限——作为自身安全评估器。该框架包含四个阶段:识别不安全工作流、在真实环境中验证、生成安全约束、验证约束有效性。输出包括包含不安全流程、测试案例和已验证约束的评估报告,支持多种安全应用。通过实证实验,我们展示了AgentGuard的可行性。本探索性工作旨在推动建立标准化的LLM智能体测试与加固流程,提升其在真实应用中的可信度。
原文摘要 · Abstract (English)
The integration of tool use into large language models (LLMs) enables agentic systems with real-world impact. In the meantime, unlike standalone LLMs, compromised agents can execute malicious workflows with more consequential impact, signified by their tool-use capability. We propose AgentGuard, a framework to autonomously discover and validate unsafe tool-use workflows, followed by generating safety constraints to confine the behaviors of agents, achieving the baseline of safety guarantee at deployment. AgentGuard leverages the LLM orchestrator's innate capabilities - knowledge of tool functionalities, scalable and realistic workflow generation, and tool execution privileges - to act as its own safety evaluator. The framework operates through four phases: identifying unsafe workflows, validating them in real-world execution, generating safety constraints, and validating constraint efficacy. The output, an evaluation report with unsafe workflows, test cases, and validated constraints, enables multiple security applications. We empirically demonstrate AgentGuard's feasibility with experiments. With this exploratory work, we hope to inspire the establishment of standardized testing and hardening procedures for LLM agents to enhance their trustworthiness in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。