提出新型多智能体安全评估框架,揭示复杂交互中的隐藏风险。
Kaleidoscopic Teaming in Multi Agent Simulations
- 设计'万花筒式协作'框架,模拟真实社会中多智能体复杂互动。
- 在单/多智能体场景下识别出多个模型的安全漏洞,验证框架有效性。
- 适合关注AI安全、智能体系统评估的研究者与开发者使用。
AI智能体因自主使用工具的能力而受到广泛关注,其自主性也带来了单智能体与多智能体场景下的新安全挑战。现有红队测试或安全评估框架难以有效评估智能体在复杂行为、思维过程及交互中暴露的风险,尤其在多智能体环境中,多种漏洞可能因协作或竞争行为被激发。为此,本文提出‘万花筒式协作’(kaleidoscopic teaming)概念,以捕捉单/多智能体场景中广泛存在的安全隐患。我们构建了一个新的评估框架,通过生成多样化的现实社会场景,对智能体在单智能体和多智能体设置下的安全性进行评测。在单智能体场景中,智能体需利用可用工具完成任务;在多智能体场景中,智能体通过合作或竞争完成任务,从而暴露潜在安全缺陷。我们引入新的上下文优化技术以生成更有效的测试场景,并设计了相应的评估指标。实验表明,该框架能有效识别多个模型在代理应用场景中的安全弱点。
原文摘要 · Abstract (English)
Warning: This paper contains content that may be inappropriate or offensive. AI agents have gained significant recent attention due to their autonomous tool usage capabilities and their integration in various real-world applications. This autonomy poses novel challenges for the safety of such systems, both in single- and multi-agent scenarios. We argue that existing red teaming or safety evaluation frameworks fall short in evaluating safety risks in complex behaviors, thought processes and actions taken by agents. Moreover, they fail to consider risks in multi-agent setups where various vulnerabilities can be exposed when agents engage in complex behaviors and interactions with each other. To address this shortcoming, we introduce the term kaleidoscopic teaming which seeks to capture complex and wide range of vulnerabilities that can happen in agents both in single-agent and multi-agent scenarios. We also present a new kaleidoscopic teaming framework that generates a diverse array of scenarios modeling real-world human societies. Our framework evaluates safety of agents in both single-agent and multi-agent setups. In single-agent setup, an agent is given a scenario that it needs to complete using the tools it has access to. In multi-agent setup, multiple agents either compete against or cooperate together to complete a task in the scenario through which we capture existing safety vulnerabilities in agents. We introduce new in-context optimization techniques that can be used in our kaleidoscopic teaming framework to generate better scenarios for safety analysis. Lastly, we present appropriate metrics that can be used along with our framework to measure safety of agents. Utilizing our kaleidoscopic teaming framework, we identify vulnerabilities in various models with respect to their safety in agentic use-cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。