arXiv:2604.24477cs.CRcs.AI2026-04被引 2

构建用于LLM多智能体系统异常监测的统一评估框架。

GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

论文配图:GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems
图 1 · 摘自论文原文
  • 设计双阶段流程:生成模拟交互数据与动态测试防御模型。
  • 在MMLU-Pro等任务上验证框架高效性,支持多种网络拓扑。
  • 适合研究安全防御或系统优化的开发者与研究人员。

大型语言模型(LLMs)在多智能体系统(MAS)中的快速集成显著提升了协作求解能力,但也扩大了攻击面,暴露于提示污染和通信被篡改等风险。尽管基于图的异常检测方法显示出潜力,但该领域仍缺乏标准化、可复现的训练与评估环境。为此,我们提出Gammaf(Graph-based Anomaly Monitoring for LLM Multi-Agent systems Framework),一个开源基准平台。Gammaf并非新型防御机制,而是一个综合性评估架构,用于生成合成多智能体交互数据集,并评测现有及未来防御模型性能。其包含两个相互依赖的流程:训练数据生成阶段,通过模拟不同网络拓扑下的辩论过程,将交互建模为鲁棒的属性图;防御系统评测阶段,在实时推理中动态隔离被标记的恶意节点,主动评估防御模型表现。通过在多个知识任务(如MMLU-Pro和GSM8K)上使用经典基线(XG-Guard与BlindGuard)进行严格评估,证明了Gammaf具有高实用性、拓扑可扩展性与执行效率。实验还表明,部署有效攻击修复机制不仅能恢复系统完整性,还能通过促进早期共识,大幅减少因恶意代理引发的冗余令牌生成,从而显著降低整体运行成本。

原文摘要 · Abstract (English)

The rapid integration of Large Language Models (LLMs) into Multi-Agent Systems (MAS) has significantly enhanced their collaborative problem-solving capabilities, but it has also expanded their attack surfaces, exposing them to vulnerabilities such as prompt infection and compromised inter-agent communication. While emerging graph-based anomaly detection methods show promise in protecting these networks, the field currently lacks a standardized, reproducible environment to train these models and evaluate their efficacy. To address this gap, we introduce Gammaf (Graph-based Anomaly Monitoring for LLM Multi-Agent systems Framework), an open-source benchmarking platform. Gammaf is not a novel defense mechanism itself, but rather a comprehensive evaluation architecture designed to generate synthetic multi-agent interaction datasets and benchmark the performance of existing and future defense models. The proposed framework operates through two interdependent pipelines: a Training Data Generation stage, which simulates debates across varied network topologies to capture interactions as robust attributed graphs, and a Defense System Benchmarking stage, which actively evaluates defense models by dynamically isolating flagged adversarial nodes during live inference rounds. Through rigorous evaluation using established defense baselines (XG-Guard and BlindGuard) across multiple knowledge tasks (such as MMLU-Pro and GSM8K), we demonstrate Gammaf's high utility, topological scalability, and execution efficiency. Furthermore, our experimental results reveal that equipping an LLM-MAS with effective attack remediation not only recovers system integrity but also substantially reduces overall operational costs by facilitating early consensus and cutting off the extensive token generation typical of adversarial agents.

多智能体异常检测安全评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。