用攻击图系统评估大模型多智能体系统的安全风险。
ATAG: AI-Agent Application Threat Assessment with Attack Graphs
- 扩展多智能体攻击图生成工具,定制规则与事实建模智能体拓扑。
- 成功生成多步攻击路径,涵盖提示注入、过度授权等漏洞。
- 适合安全研究人员和开发团队用于预判智能体应用风险。
基于大语言模型的多智能体系统(MASs)安全评估面临挑战,主要源于系统内部动态复杂性及大模型漏洞的持续演化。传统攻击图(AG)方法缺乏对大模型攻击的建模能力。本文提出AI-agent应用威胁评估框架ATAG,通过扩展基于逻辑的MulVAL攻击图生成工具,引入自定义事实与交互规则,准确刻画智能体拓扑、漏洞与攻击场景。研究还构建了大模型漏洞数据库(LVD),推动漏洞文档标准化。在两个多智能体应用上验证显示,ATAG能有效建模并生成利用提示注入、过度代理、敏感信息泄露及不安全输出处理等漏洞的复杂多步攻击路径。该框架为理解、可视化和优先排序多智能体人工智能系统(MAASs)中的复杂攻击路径提供了重要工具,助力主动识别与缓解智能体应用威胁。
原文摘要 · Abstract (English)
Evaluating the security of multi-agent systems (MASs) powered by large language models (LLMs) is challenging, primarily because of the systems' complex internal dynamics and the evolving nature of LLM vulnerabilities. Traditional attack graph (AG) methods often lack the specific capabilities to model attacks on LLMs. This paper introduces AI-agent application Threat assessment with Attack Graphs (ATAG), a novel framework designed to systematically analyze the security risks associated with AI-agent applications. ATAG extends the MulVAL logic-based AG generation tool with custom facts and interaction rules to accurately represent AI-agent topologies, vulnerabilities, and attack scenarios. As part of this research, we also created the LLM vulnerability database (LVD) to initiate the process of standardizing LLM vulnerabilities documentation. To demonstrate ATAG's efficacy, we applied it to two multi-agent applications. Our case studies demonstrated the framework's ability to model and generate AGs for sophisticated, multi-step attack scenarios exploiting vulnerabilities such as prompt injection, excessive agency, sensitive information disclosure, and insecure output handling across interconnected agents. ATAG is an important step toward a robust methodology and toolset to help understand, visualize, and prioritize complex attack paths in multi-agent AI systems (MAASs). It facilitates proactive identification and mitigation of AI-agent threats in multi-agent applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。