通过相互推理检测多智能体中的后门攻击,提升系统安全性。
PeerGuard: Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning
- 让智能体互评推理逻辑,识别异常响应以发现被污染的代理。
- 在ChatGPT和Llama 3系统中实现高精度毒化代理识别,误报率低。
- 适合关注多智能体安全与可信AI交互的研究者与开发者。
多智能体系统利用先进AI模型作为自主代理,在机器人、交通管理等场景中协同或竞争完成复杂任务。尽管其重要性日益凸显,多智能体系统的安全性仍鲜受关注,多数研究集中于单个AI模型而非交互式代理。本文探究多智能体系统中的后门漏洞,并提出基于代理间互动的防御机制。通过利用推理能力,每个代理评估其他代理的响应,以检测不合逻辑的推理过程,从而识别中毒代理。在基于LLM的多智能体系统(包括ChatGPT系列和Llama 3)上的实验表明,该方法能有效识别中毒代理,同时对干净代理的误报率极低。本工作为多智能体系统安全提供了新视角,助力构建更鲁棒、可信的AI交互体系。
原文摘要 · Abstract (English)
Multi-agent systems leverage advanced AI models as autonomous agents that interact, cooperate, or compete to complete complex tasks across applications such as robotics and traffic management. Despite their growing importance, safety in multi-agent systems remains largely underexplored, with most research focusing on single AI models rather than interacting agents. This work investigates backdoor vulnerabilities in multi-agent systems and proposes a defense mechanism based on agent interactions. By leveraging reasoning abilities, each agent evaluates responses from others to detect illogical reasoning processes, which indicate poisoned agents. Experiments on LLM-based multi-agent systems, including ChatGPT series and Llama 3, demonstrate the effectiveness of the proposed method, achieving high accuracy in identifying poisoned agents while minimizing false positives on clean agents. We believe this work provides insights into multi-agent system safety and contributes to the development of robust, trustworthy AI interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。