用随机平滑提升大模型多智能体系统的抗攻击能力。
Enhancing Robustness of LLM-Driven Multi-Agent Systems through Randomized Smoothing
- 将随机平滑用于多智能体共识,实现对抗干扰下的概率安全保证。
- 在仿真中有效阻止对抗行为与幻觉传播,同时保持共识性能。
- 适合高风险场景下大模型多智能体系统的安全部署。
本文提出一种防御框架,用于增强大语言模型(LLM)驱动的多智能体系统(MAS)在航空航天等安全关键领域中的安全性。我们应用随机平滑——一种统计鲁棒性认证技术——于多智能体共识场景,可在对抗干扰下提供决策的概率保证。与传统验证方法不同,本方法在黑盒设置下运行,并采用两阶段自适应采样机制,在鲁棒性与计算效率间取得平衡。仿真结果表明,该方法能有效防止对抗行为和幻觉的传播,同时维持共识性能。本工作为大模型驱动的多智能体系统在真实高风险环境中的安全部署提供了可行且可扩展的路径。
原文摘要 · Abstract (English)
This paper presents a defense framework for enhancing the safety of large language model (LLM) empowered multi-agent systems (MAS) in safety-critical domains such as aerospace. We apply randomized smoothing, a statistical robustness certification technique, to the MAS consensus context, enabling probabilistic guarantees on agent decisions under adversarial influence. Unlike traditional verification methods, our approach operates in black-box settings and employs a two-stage adaptive sampling mechanism to balance robustness and computational efficiency. Simulation results demonstrate that our method effectively prevents the propagation of adversarial behaviors and hallucinations while maintaining consensus performance. This work provides a practical and scalable path toward safe deployment of LLM-based MAS in real-world, high-stakes environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。