arXiv:2603.15661cs.AI2026-03被引 2

动态信任图防御多智能体系统中的潜伏攻击者

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

  • 构建动态信任图,实时更新智能体信任度
  • 防御成功率超86%,误报率显著降低
  • 适合高安全性要求的协作式AI系统

基于大语言模型的多智能体系统展现出强大的协同推理能力,但也引入了新型攻击面,如潜伏攻击者:在常规运行中表现正常,逐步积累信任,仅在特定条件触发后才暴露恶意行为。现有防御方法多依赖静态图优化或层级数据管理,难以应对不断演化的攻击策略,且因僵化封禁政策导致误报率过高。为此,我们提出DynaTrust,一种针对潜伏攻击者的新型防御方法。DynaTrust将多智能体系统建模为动态信任图(DTG),将信任视为连续演化过程而非静态属性,依据历史行为和专家代理的置信度动态更新各智能体信任值。不简单封禁,而是自主重构图结构以隔离受损智能体并恢复任务连通性,保障系统可用性。我们在基于AdvBench和HumanEval的混合基准上评估DynaTrust,结果表明其相较最优方法AgentShield提升41.7%的防御成功率,在对抗条件下达到超过86%的防御率;同时显著降低误报率,通过图适配确保系统持续稳定运行。

原文摘要 · Abstract (English)

Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable collaborative reasoning capabilities but introduce new attack surfaces, such as the sleeper agent, which behave benignly during routine operation and gradually accumulate trust, only revealing malicious behaviors when specific conditions or triggers are met. Existing defense works primarily focus on static graph optimization or hierarchical data management, often failing to adapt to evolving adversarial strategies or suffering from high false-positive rates (FPR) due to rigid blocking policies. To address this, we propose DynaTrust, a novel defense method against sleeper agents. DynaTrust models MAS as a dynamic trust graph~(DTG), and treats trust as a continuous, evolving process rather than a static attribute. It dynamically updates the trust of each agent based on its historical behaviors and the confidence of selected expert agents. Instead of simply blocking, DynaTrust autonomously restructures the graph to isolate compromised agents and restore task connectivity to ensure the usability of MAS. To assess the effectiveness of DynaTrust, we evaluate it on mixed benchmarks derived from AdvBench and HumanEval. The results demonstrate that DynaTrust outperforms the state-of-the-art method AgentShield by increasing the defense success rate by 41.7%, achieving rates exceeding 86% under adversarial conditions. Furthermore, it effectively balances security with utility by significantly reducing FPR, ensuring uninterrupted system operations through graph adaptation.

多智能体安全防御动态信任大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。