通过伪装通信内容,隐蔽攻击大模型多智能体系统
Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS
- 用树搜索与偏好优化训练攻击策略,动态生成多轮干扰内容
- 在多种任务和模型上攻击成功率超基线,且更难被发现
- 适合研究系统安全或对抗攻防的学者参考
基于大语言模型的多智能体系统(LLM-MAS)通过智能体间通信完成复杂动态任务,但这种依赖带来显著安全风险。现有攻击方法或针对智能体内部,或依赖直接显性说服,限制了其有效性、适应性和隐蔽性。本文提出MAST框架,一种多轮自适应隐蔽篡改方法,旨在利用系统通信漏洞。MAST结合蒙特卡洛树搜索与直接偏好优化,训练攻击策略模型,以自适应生成高效的多轮篡改方案。为保持隐蔽性,篡改过程引入语义与嵌入相似性双重约束。在多种任务、通信架构及大模型上的全面实验表明,MAST在攻击成功率上持续优于基线,同时显著提升隐蔽性。结果凸显了MAST的有效性、隐蔽性与适应性,强调了在LLM-MAS中建立强通信防护机制的必要性。
原文摘要 · Abstract (English)
Large language model-based multi-agent systems (LLM-MAS) effectively accomplish complex and dynamic tasks through inter-agent communication, but this reliance introduces substantial safety vulnerabilities. Existing attack methods targeting LLM-MAS either compromise agent internals or rely on direct and overt persuasion, which limit their effectiveness, adaptability, and stealthiness. In this paper, we propose MAST, a Multi-round Adaptive Stealthy Tampering framework designed to exploit communication vulnerabilities within the system. MAST integrates Monte Carlo Tree Search with Direct Preference Optimization to train an attack policy model that adaptively generates effective multi-round tampering strategies. Furthermore, to preserve stealthiness, we impose dual semantic and embedding similarity constraints during the tampering process. Comprehensive experiments across diverse tasks, communication architectures, and LLMs demonstrate that MAST consistently achieves high attack success rates while significantly enhancing stealthiness compared to baselines. These findings highlight the effectiveness, stealthiness, and adaptability of MAST, underscoring the need for robust communication safeguards in LLM-MAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。