arXiv:2505.12442cs.CRcs.AI2025-05被引 8

攻击者可黑盒窃取多智能体系统核心机密,成功率超87%。

IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems

  • 设计黑盒攻击框架MASLEAK,通过精心构造查询触发信息泄露。
  • 在真实系统上实现87%成功率提取系统提示与任务指令,92%还原架构。
  • 适用于评估多智能体系统安全,提醒开发者加强隐私防护。

大型语言模型的快速发展催生了通过协作完成复杂任务的多智能体系统(MAS)。然而,其复杂的架构与智能体交互机制带来了显著的知识产权保护风险。本文提出MASLEAK,一种新型黑盒攻击框架,旨在从MAS应用中提取敏感信息。攻击者仅可通过公开API提交查询$ q $并观察最终智能体输出,无须了解系统架构或配置。受计算机蠕虫传播启发,MASLEAK精心设计对抗性查询,诱导、传播并保留各智能体的响应,以获取完整专有组件,包括智能体数量、系统拓扑、系统提示、任务指令和工具使用情况。我们构建了首个包含810个应用的合成数据集,并在实际系统(如Coze和CrewAI)上评估了该攻击。结果显示,系统提示与任务指令的平均攻击成功率高达87%,系统架构的提取成功率在多数情况下达到92%。最后讨论了研究影响及潜在防御策略。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has led to the emergence of Multi-Agent Systems (MAS) to perform complex tasks through collaboration. However, the intricate nature of MAS, including their architecture and agent interactions, raises significant concerns regarding intellectual property (IP) protection. In this paper, we introduce MASLEAK, a novel attack framework designed to extract sensitive information from MAS applications. MASLEAK targets a practical, black-box setting, where the adversary has no prior knowledge of the MAS architecture or agent configurations. The adversary can only interact with the MAS through its public API, submitting attack query $q$ and observing outputs from the final agent. Inspired by how computer worms propagate and infect vulnerable network hosts, MASLEAK carefully crafts adversarial query $q$ to elicit, propagate, and retain responses from each MAS agent that reveal a full set of proprietary components, including the number of agents, system topology, system prompts, task instructions, and tool usages. We construct the first synthetic dataset of MAS applications with 810 applications and also evaluate MASLEAK against real-world MAS applications, including Coze and CrewAI. MASLEAK achieves high accuracy in extracting MAS IP, with an average attack success rate of 87% for system prompts and task instructions, and 92% for system architecture in most cases. We conclude by discussing the implications of our findings and the potential defenses.

多智能体安全攻防隐私泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。