用强化学习让攻击者自进化,突破工具攻击的局限性
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS

- 构建动态记忆库,通过反思机制生成对抗性策略
- 在多智能体系统中实现长时程攻击,成功率显著超越基线
- 适合研究模型安全与防御机制的研究者参考
基于大语言模型的多智能体系统(LLM-MAS)通过协调专用智能体和外部工具解决复杂任务,但对工具输出的隐式信任带来了关键攻击面。现有工具攻击受限于领域特定性或固定模板。为此,我们提出Evo-Attacker,将工具攻击建模为自演化、带记忆增强的强化学习过程。该方法构建动态攻击记忆,利用深思熟虑的推理机制,在关键时刻检索对抗模式并制定修改策略。此外,引入Attack-Flow GRPO,通过最终结果优化中间推理步骤,解决长时程信用分配难题。全面实验表明,Evo-Attacker持续优于基线,展现出优异的泛化与演化能力,凸显防御工具安全的紧迫性。
原文摘要 · Abstract (English)
While Large Language Model-based Multi-Agent Systems (LLM-MAS) demonstrate remarkable capabilities in solving complex tasks by orchestrating specialized agents and external tools, the implicit trust in tool outputs creates a critical attack surface. Existing tool attacks are limited by domain specificity or fixed and static templates. To address these challenges, we propose Evo-Attacker, which formulates the tool attack as a self-evolving, memory-augmented reinforcement learning process. Evo-Attacker constructs a dynamic attack memory and employs deliberative reasoning to retrieve adversarial patterns and strategize modifying interventions at critical moments. Furthermore, we introduce Attack-Flow GRPO to optimize intermediate reasoning steps via terminal outcomes, addressing the long-horizon credit assignment challenge. Comprehensive experiments demonstrate that Evo-Attacker consistently outperforms baselines, highlighting its generalization and evolutionary capabilities and the urgent need for defensive tool safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。