arXiv:2601.14667cs.MAcs.AI2026-01被引 4

提出新型防御框架,识别并修复被攻陷的智能体,防止恶意信息扩散。

INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems

  • 通过感染感知检测与拓扑约束定位攻击源和受影响范围。
  • 平均降低33%攻击成功率,且在多种模型间表现稳定。
  • 适合关注多智能体系统安全的开发者与研究人员。

基于大语言模型的多智能体系统快速发展,但恶意影响可通过智能体间通信呈病毒式传播。传统防护方法采用二元区分机制,难以识别被攻击者劫持的良性智能体。本文提出感染感知防护框架INFA-Guard,将被感染智能体作为独立威胁类别进行处理。该框架利用感染感知检测与拓扑约束,精准定位攻击源及感染范围;在修复阶段,替换攻击者并恢复被感染智能体,避免恶意传播的同时保持系统拓扑完整性。大量实验表明,INFA-Guard达到当前最优性能,平均降低攻击成功率(ASR)33%,具备跨模型鲁棒性、优越的拓扑泛化能力及高成本效益。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Model (LLM)-based Multi-Agent Systems (MAS) has introduced significant security vulnerabilities, where malicious influence can propagate virally through inter-agent communication. Conventional safeguards often rely on a binary paradigm that strictly distinguishes between benign and attack agents, failing to account for infected agents i.e., benign entities converted by attack agents. In this paper, we propose Infection-Aware Guard, INFA-Guard, a novel defense framework that explicitly identifies and addresses infected agents as a distinct threat category. By leveraging infection-aware detection and topological constraints, INFA-Guard accurately localizes attack sources and infected ranges. During remediation, INFA-Guard replaces attackers and rehabilitates infected ones, avoiding malicious propagation while preserving topological integrity. Extensive experiments demonstrate that INFA-Guard achieves state-of-the-art performance, reducing the Attack Success Rate (ASR) by an average of 33%, while exhibiting cross-model robustness, superior topological generalization, and high cost-effectiveness.

多智能体安全防护大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。