通过双层图检测,精准识别大模型多智能体中的恶意行为并解释原因。
Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
- 双层编码器融合句子与词汇级信息,提升检测精度。
- 在多种拓扑和攻击场景下,检测准确率显著优于基线方法。
- 可解释性强,能定位异常行为的具体词语,适合安全审查场景。
基于大语言模型(LLM)的多智能体系统(MAS)在解决复杂任务中展现出强大能力。随着其在诸多安全关键任务中日益自主,识别恶意智能体已成为关键安全挑战。现有基于图异常检测(GAD)的防御方法主要依赖粗粒度的句子级信息,忽视细粒度的词汇线索,导致性能受限。同时,这些方法缺乏可解释性,限制了其可靠性与实际应用。为此,我们提出XG-Guard——一种可解释且细粒度的守护框架,用于检测MAS中的恶意智能体。通过双层智能体编码器,联合建模每个智能体的句子级与词元级表征;基于主题的异常检测器捕捉对话中讨论焦点的动态演变;双层得分融合机制量化词元级别的贡献以实现解释。在多种MAS拓扑结构与攻击场景下的大量实验表明,XG-Guard具有鲁棒的检测性能与强可解释性。
原文摘要 · Abstract (English)
Large language model (LLM)-based multi-agent systems (MAS) have shown strong capabilities in solving complex tasks. As MAS become increasingly autonomous in various safety-critical tasks, detecting malicious agents has become a critical security concern. Although existing graph anomaly detection (GAD)-based defenses can identify anomalous agents, they mainly rely on coarse sentence-level information and overlook fine-grained lexical cues, leading to suboptimal performance. Moreover, the lack of interpretability in these methods limits their reliability and real-world applicability. To address these limitations, we propose XG-Guard, an explainable and fine-grained safeguarding framework for detecting malicious agents in MAS. To incorporate both coarse and fine-grained textual information for anomalous agent identification, we utilize a bi-level agent encoder to jointly model the sentence- and token-level representations of each agent. A theme-based anomaly detector further captures the evolving discussion focus in MAS dialogues, while a bi-level score fusion mechanism quantifies token-level contributions for explanation. Extensive experiments across diverse MAS topologies and attack scenarios demonstrate robust detection performance and strong interpretability of XG-Guard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。