多智能体系统中,恶意提示会像病毒一样自我复制传播。
Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- 设计了可跨智能体传播的恶意提示攻击,类比计算机病毒。
- 实验证明即使不公开通信,系统仍极易被感染。
- 提出标签防御法,结合现有措施显著降低传播风险。
随着大语言模型(LLMs)能力增强,多智能体系统在现代AI应用中日益普及。然而,现有安全研究主要聚焦单智能体场景下的漏洞,如提示注入攻击——恶意提示嵌入外部内容,诱使模型执行非预期或有害操作。本文揭示了一种更危险的攻击路径:多智能体系统内的LLM-to-LLM提示注入。我们提出Prompt Infection,一种恶意提示在互联智能体间自我复制的新型攻击,行为类似计算机病毒。该攻击可能导致数据窃取、诈骗、误导信息及系统瘫痪,且在系统内静默传播。大量实验表明,即便智能体不公开全部通信,系统仍高度易受攻击。为此,我们提出LLM Tagging防御机制,与现有防护结合后能显著抑制感染扩散。本工作凸显了多智能体大模型系统普及背景下亟需强化安全机制。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) grow increasingly powerful, multi-agent systems are becoming more prevalent in modern AI applications. Most safety research, however, has focused on vulnerabilities in single-agent LLMs. These include prompt injection attacks, where malicious prompts embedded in external content trick the LLM into executing unintended or harmful actions, compromising the victim's application. In this paper, we reveal a more dangerous vector: LLM-to-LLM prompt injection within multi-agent systems. We introduce Prompt Infection, a novel attack where malicious prompts self-replicate across interconnected agents, behaving much like a computer virus. This attack poses severe threats, including data theft, scams, misinformation, and system-wide disruption, all while propagating silently through the system. Our extensive experiments demonstrate that multi-agent systems are highly susceptible, even when agents do not publicly share all communications. To address this, we propose LLM Tagging, a defense mechanism that, when combined with existing safeguards, significantly mitigates infection spread. This work underscores the urgent need for advanced security measures as multi-agent LLM systems become more widely adopted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。