arXiv:2603.15408cs.CRcs.AI2026-03被引 1

为大模型多智能体系统设计统一安全防护框架,识别20类风险并实现实时监控。

TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems

  • 构建三层次风险分类体系,覆盖单智能体、通信与系统级风险
  • 通过攻击探测生成详细漏洞报告,支持开发前评估与运行时监控
  • 适配多种多智能体结构,可扩展性强,适合研究与工程落地

随着基于大语言模型的多智能体系统(MAS)快速发展,其安全与隐私问题日益凸显,带来超越单个智能体或大模型的新风险。现有研究缺乏针对多智能体系统风险的统一防护机制。本文提出TrinityGuard,一个基于OWASP标准的综合性安全评估与监控框架。该框架包含三层细粒度风险分类,识别出20种风险类型,涵盖单智能体漏洞、智能体间通信威胁及系统级涌现危害。TrinityGuard采用三位一体设计:智能体抽象层可适配各类多智能体结构,评估层包含针对性测试模块,运行时监控代理由统一的LLM评判工厂协调。评估阶段,框架执行定制化攻击探测,生成各风险类型的详细漏洞报告;监控代理分析结构化执行轨迹并实时预警,支持开发前评估与运行时监控。我们进一步形式化安全指标,并在多个典型多智能体案例中展示其通用性与可靠性。TrinityGuard为多智能体系统的风险评估与监控提供了全面解决方案,推动该领域安全研究发展。

原文摘要 · Abstract (English)

With the rapid development of LLM-based multi-agent systems (MAS), their significant safety and security concerns have emerged, which introduce novel risks going beyond single agents or LLMs. Despite attempts to address these issues, the existing literature lacks a cohesive safeguarding system specialized for MAS risks. In this work, we introduce TrinityGuard, a comprehensive safety evaluation and monitoring framework for LLM-based MAS, grounded in the OWASP standards. Specifically, TrinityGuard encompasses a three-tier fine-grained risk taxonomy that identifies 20 risk types, covering single-agent vulnerabilities, inter-agent communication threats, and system-level emergent hazards. Designed for scalability across various MAS structures and platforms, TrinityGuard is organized in a trinity manner, involving an MAS abstraction layer that can be adapted to any MAS structures, an evaluation layer containing risk-specific test modules, alongside runtime monitor agents coordinated by a unified LLM Judge Factory. During Evaluation, TrinityGuard executes curated attack probes to generate detailed vulnerability reports for each risk type, where monitor agents analyze structured execution traces and issue real-time alerts, enabling both pre-development evaluation and runtime monitoring. We further formalize these safety metrics and present detailed case studies across various representative MAS examples, showcasing the versatility and reliability of TrinityGuard. Overall, TrinityGuard acts as a comprehensive framework for evaluating and monitoring various risks in MAS, paving the way for further research into their safety and security.

多智能体安全评估大模型风险监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。