arXiv:2605.28104cs.AI2026-05被引 2

提出防御协作攻击的句子级修正方法,提升多智能体系统可靠性

Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification

论文配图:Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification
图 1 · 摘自论文原文
  • 设计自适应协作攻击框架,恶意智能体通过多轮交互协同误导
  • 协作攻击使任务成功率下降5.34%,而新方法平均提升36.76%
  • 在句子级别识别并修正误导信息,适合安全敏感的多智能体场景

近年来,基于大语言模型的多智能体系统(MAS)在协同决策与复杂问题求解方面表现出色。然而,恶意智能体可能注入虚假信息以误导其他智能体并破坏系统性能,催生了针对攻击机制与防御策略的新研究方向。以往研究多假设恶意智能体独立行动,但本文指出其可能呈现协作行为,通过内部信息交换实现更高效的攻击。为此,我们提出一种自适应协作攻击框架,使恶意智能体能自主协调并动态调整攻击策略。同时,提出句子级可信度分析与修正(STAR)防御框架,可在智能体通信中识别并修正误导性语句。实验表明,协作攻击导致任务成功率显著下降,相对降幅达5.34%;而STAR能有效缓解协作与独立攻击,平均提升任务成功率36.76%。代码已公开于https://github.com/smoooom/STAR。

原文摘要 · Abstract (English)

Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving. However, malicious agents in MAS may inject misinformation to mislead other agents and disrupt system performance, giving rise to a new research direction that focuses on attack mechanisms and defense strategies in MAS. Prior studies largely assume malicious agents act independently and investigate the corresponding defense strategies. However, we argue that malicious agents may exhibit collaborative behaviors, enabling more effective attacks through internal information exchange. In this paper, we propose an adaptive cooperative attack framework, where malicious agents autonomously coordinate and dynamically adjust their attack strategies through multi-round interactions. Furthermore, we introduce Sentence-Level Trustworthiness Analysis and Rectification (STAR), a defense framework that identifies and rectifies misleading information at the sentence level within agent communications. Our experiments show that cooperative attacks lead to a significantly larger degradation in task success rate than independent attacks, resulting in a relative drop of 5.34\%. Meanwhile, STAR effectively mitigates both cooperative and independent threats and improves task success rate by an average of 36.76\%. The code is available at https://github.com/smoooom/STAR.

多智能体协同攻击防御机制大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。