多智能体系统中,证据阈值触发后门攻击,新防御可有效拦截。
When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems

- 通过隐藏阈值触发集体后门,仅当多方证据累积达标时激活。
- 攻击在不暴露目标和触发条件的情况下实现精准激活,干扰极小。
- 新防御机制检测异常通信动态,防止恶意更新扩散。
基于大模型的多智能体系统(MAS)通过迭代沟通与共享上下文扩展能力,但协作也引入了新漏洞:后门行为可在同伴证据达到隐藏阈值时被触发,而非由单条消息决定。本文提出一种集体证据阈值后门范式及边界条件后门注入方法(BCBI),通过构建反事实边界对,区分阈值前的正常行为与阈值后的恶意目标,并学习与证据进展一致的潜在状态演化。为应对该威胁,提出干净样本潜变量转移测试时评估(LATTE)防御机制,通过学习良性通信动态,在异常智能体更新传播前将其隔离。在多个基准测试中,BCBI实现了选择性激活且提前误触发极少;即便不知攻击目标或触发条件,LATTE也能有效限制传播,干扰最小。
原文摘要 · Abstract (English)
LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. We introduce a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection (BCBI), which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence. To mitigate this threat, we propose LAtent Transition Test-time Evaluation (LATTE), a clean-only latent-transition defense that learns benign communication dynamics and quarantines anomalous agent updates before their responses propagate. Across several benchmarks, BCBI yields selective activation with little premature activation; without knowing the attack target or trigger, LATTE limits propagation with minimal disruption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。