无需控制管理员,仅用单个代理即可隐蔽推断大模型多智能体系统拓扑结构
WebWeaver: Breaking Topology Confidentiality in LLM Multi-Agent Systems with Stealthy Context-Based Inference
- 通过分析代理上下文而非身份,实现对系统拓扑的隐蔽推断
- 在主动防御下仍达60%更高准确率,且计算开销极低
- 适用于研究系统安全或对抗攻防的从业者
通信拓扑是大语言模型多智能体系统(LLM-MAS)实用性和安全性的重要因素,被视为高价值知识产权,但其保密性尚未得到充分研究。现有拓扑推断方法依赖不切实际的假设,如控制管理代理或通过越狱直接查询身份,极易被基础关键词防御手段挫败。因此,先前分析未能反映真实威胁。为弥合这一现实差距,我们提出WebWeaver攻击框架,仅需攻破任意一个普通代理即可推断完整拓扑结构。与以往方法不同,WebWeaver仅依赖代理上下文而非身份标识,实现更强隐蔽性。该框架进一步引入新型隐秘越狱机制和全新全无越狱的扩散设计,以应对越狱失败场景。此外,针对扩散推断中的关键挑战,提出一种掩码策略,在扩散过程中保留已知拓扑信息,并提供正确性理论保障。大量实验表明,WebWeaver显著优于当前最优基线,在主动防御下推理准确率高出约60%,且开销可忽略。
原文摘要 · Abstract (English)
Communication topology is a critical factor in the utility and safety of LLM-based multi-agent systems (LLM-MAS), making it a high-value intellectual property (IP) whose confidentiality remains insufficiently studied. Existing topology inference attempts rely on impractical assumptions, including control over the administrative agent and direct identity queries via jailbreaks, which are easily defeated by basic keyword-based defenses. As a result, prior analyses fail to capture the real-world threat of such attacks. To bridge this realism gap, we propose \textit{WebWeaver}, an attack framework that infers the complete LLM-MAS topology by compromising only a single arbitrary agent instead of the administrative agent. Unlike prior approaches, WebWeaver relies solely on agent contexts rather than agent IDs, enabling significantly stealthier inference. WebWeaver further introduces a new covert jailbreak-based mechanism and a novel fully jailbreak-free diffusion design to handle cases where jailbreaks fail. Additionally, we address a key challenge in diffusion-based inference by proposing a masking strategy that preserves known topology during diffusion, with theoretical guarantees of correctness. Extensive experiments show that WebWeaver substantially outperforms state-of-the-art (SOTA) baselines, achieving about 60\% higher inference accuracy under active defenses with negligible overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。