构建可定制的多智能体安全测试平台,研究大模型协作中的隐私与安全风险。
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
- 复用早期黑板架构,搭建模块化多智能体实验环境。
- 识别出误对齐、恶意代理、通信劫持等四类关键攻击向量。
- 支持快速验证防御方案,适合安全研究者和系统开发者使用。
基于大语言模型的多智能体系统(MAS)可自动化处理会议安排等需协作的任务,利用大模型实现对非结构化私有数据、用户约束和偏好的细致处理。然而,这种设计引入了新型风险,包括对齐错误及恶意方攻击导致的代理被控制或用户数据泄露。本文提出Terrarium框架,用于细粒度研究基于大模型的多智能体系统在安全、隐私与安全方面的挑战。通过重用多智能体系统早期的黑板设计,构建了一个模块化、可配置的测试平台,以支持多智能体协作研究。我们识别出误对齐、恶意代理、通信通道被攻破、数据投毒等关键攻击向量,并实现了三个协作场景与四种代表性攻击,验证了框架的灵活性。通过提供快速原型设计、评估与迭代防御机制的工具,Terrarium旨在加速可信多智能体系统的进展。
原文摘要 · Abstract (English)
A multi-agent system (MAS) powered by large language models (LLMs) can automate tedious user tasks such as meeting scheduling that requires inter-agent collaboration. LLMs enable nuanced protocols that account for unstructured private data, user constraints, and preferences. However, this design introduces new risks, including misalignment and attacks by malicious parties that compromise agents or steal user data. In this paper, we propose the Terrarium framework for fine-grained study on safety, privacy, and security in LLM-based MAS. We repurpose the blackboard design, an early approach in multi-agent systems, to create a modular, configurable testbed for multi-agent collaboration. We identify key attack vectors such as misalignment, malicious agents, compromised communication, and data poisoning. We implement three collaborative MAS scenarios with four representative attacks to demonstrate the framework's flexibility. By providing tools to rapidly prototype, evaluate, and iterate on defenses and designs, Terrarium aims to accelerate progress toward trustworthy multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。