为可控环境下大模型多智能体系统设计风险分析工具
Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems
- 提出六类关键故障模式及对应分析工具
- 通过分阶段测试逐步验证系统安全性
- 适合组织级AI治理与安全团队使用
组织正逐步采用基于大模型的AI智能体,其部署从单个智能体演变为互联的多智能体网络。然而,单个智能体安全并不等于整体安全,因智能体间随时间交互会引发涌现行为并导致新型故障模式。因此,多智能体系统需采用与单智能体截然不同的风险分析方法。本报告聚焦于受控环境中多智能体系统的早期风险识别与分析,考察六类关键故障模式:级联可靠性失效、智能体间通信失败、单一化崩溃、从众偏差、心智理论不足以及混合动机动态。针对每种模式,提供可集成到现有框架中的实践工具包。鉴于当前对大模型行为理解的根本局限,方法强调分析有效性,倡导通过逐步抽象和部署阶段的分步测试,渐进式提升暴露度并收集模拟、观测、基准测试与红队测试等多方证据。该方法为组织在部署和运营此类系统时建立稳健的风险管理基础。
原文摘要 · Abstract (English)
Organisations are starting to adopt LLM-based AI agents, with their deployments naturally evolving from single agents towards interconnected, multi-agent networks. Yet a collection of safe agents does not guarantee a safe collection of agents, as interactions between agents over time create emergent behaviours and induce novel failure modes. This means multi-agent systems require a fundamentally different risk analysis approach than that used for a single agent. This report addresses the early stages of risk identification and analysis for multi-agent AI systems operating within governed environments where organisations control their agent configurations and deployment. In this setting, we examine six critical failure modes: cascading reliability failures, inter-agent communication failures, monoculture collapse, conformity bias, deficient theory of mind, and mixed motive dynamics. For each, we provide a toolkit for practitioners to extend or integrate into their existing frameworks to assess these failure modes within their organisational contexts. Given fundamental limitations in current LLM behavioural understanding, our approach centres on analysis validity, and advocates for progressively increasing validity through staged testing across stages of abstraction and deployment that gradually increases exposure to potential negative impacts, while collecting convergent evidence through simulation, observational analysis, benchmarking, and red teaming. This methodology establishes the groundwork for robust organisational risk management as these LLM-based multi-agent systems are deployed and operated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。