提出防御多智能体系统中谣言传播的新框架,提升系统抗误导能力。
Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems
- 基于目标感知推理的两阶段无训练防御机制
- 平均降低28.17%谣言毒性,攻击下任务成功率提升10.33%
- 适用于增强复杂任务中AI系统的可信度与鲁棒性
基于大语言模型的多智能体系统在解决复杂现实任务中表现出显著优势,但因其引入了额外攻击面,易受误导信息注入影响。为深入理解此类系统中的误导信息传播动态,我们构建了MisinfoTask数据集,包含复杂且贴近现实的任务,用于评估多智能体系统的鲁棒性。在此基础上,提出ARGUS框架——一种两阶段、无需训练的目标感知防御机制,可精准识别并修正信息流中的误导内容。实验表明,在多种攻击场景下,ARGUS能有效降低误导信息毒性,平均减少约28.17%,并在攻击条件下将任务成功率提升约10.33%。代码与数据集已公开于https://github.com/zhrli324/ARGUS。
原文摘要 · Abstract (English)
Large Language Model-based Multi-Agent Systems (MASs) have demonstrated strong advantages in addressing complex real-world tasks. However, due to the introduction of additional attack surfaces, MASs are particularly vulnerable to misinformation injection. To facilitate a deeper understanding of misinformation propagation dynamics within these systems, we introduce MisinfoTask, a novel dataset featuring complex, realistic tasks designed to evaluate MAS robustness against such threats. Building upon this, we propose ARGUS, a two-stage, training-free defense framework leveraging goal-aware reasoning for precise misinformation rectification within information flows. Our experiments demonstrate that in challenging misinformation scenarios, ARGUS exhibits significant efficacy across various injection attacks, achieving an average reduction in misinformation toxicity of approximately 28.17% and improving task success rates under attack by approximately 10.33%. Our code and dataset is available at: https://github.com/zhrli324/ARGUS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。