用哨兵代理监控多智能体系统,防范各类安全威胁。
Sentinel Agents for Secure and Trustworthy Agentic AI in Multi-Agent Systems
- 部署哨兵代理网络,通过大模型语义分析与行为检测实现分布式防护。
- 在模拟中成功识别162次攻击,包括提示注入和数据泄露。
- 适合需要高安全性和合规性的智能体协作场景。
本文提出一种新型架构框架,以提升多智能体系统(MAS)的安全性与可靠性。核心是部署一组哨兵代理,作为分布式安全层,融合大型语言模型(LLM)语义分析、行为分析、检索增强验证及跨代理异常检测等技术,可实时监控智能体间通信,识别潜在威胁,强制执行隐私与访问控制,并生成完整审计记录。同时引入协调代理,负责政策执行管理与代理参与调控,并接收哨兵代理的告警,据此动态调整策略、隔离异常代理,遏制威胁扩散。该双层安全机制有效应对提示注入、共谋行为、大模型幻觉、隐私泄露及协同攻击等风险。实验在多智能体对话环境中注入162种不同类型的合成攻击,哨兵代理均成功检测到攻击行为,验证了该监控方案的可行性。框架还增强了系统可观测性,支持监管合规,并支持策略持续演进。
原文摘要 · Abstract (English)
This paper proposes a novel architectural framework aimed at enhancing security and reliability in multi-agent systems (MAS). A central component of this framework is a network of Sentinel Agents, functioning as a distributed security layer that integrates techniques such as semantic analysis via large language models (LLMs), behavioral analytics, retrieval-augmented verification, and cross-agent anomaly detection. Such agents can potentially oversee inter-agent communications, identify potential threats, enforce privacy and access controls, and maintain comprehensive audit records. Complementary to the idea of Sentinel Agents is the use of a Coordinator Agent. The Coordinator Agent supervises policy implementation, and manages agent participation. In addition, the Coordinator also ingests alerts from Sentinel Agents. Based on these alerts, it can adapt policies, isolate or quarantine misbehaving agents, and contain threats to maintain the integrity of the MAS ecosystem. This dual-layered security approach, combining the continuous monitoring of Sentinel Agents with the governance functions of Coordinator Agents, supports dynamic and adaptive defense mechanisms against a range of threats, including prompt injection, collusive agent behavior, hallucinations generated by LLMs, privacy breaches, and coordinated multi-agent attacks. In addition to the architectural design, we present a simulation study where 162 synthetic attacks of different families (prompt injection, hallucination, and data exfiltration) were injected into a multi-agent conversational environment. The Sentinel Agents successfully detected the attack attempts, confirming the practical feasibility of the proposed monitoring approach. The framework also offers enhanced system observability, supports regulatory compliance, and enables policy evolution over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。