提出多步触发后门攻击,让智能体更鲁棒却更易被操控。
Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
- 设计多步触发序列,逐步引导智能体偏离原任务。
- 攻击成功率近100%,误触发率接近0,隐蔽性强。
- 反直觉提升良性任务表现,适合研究安全与对抗攻击者。
大型语言模型驱动的智能体在实际应用中快速部署,其可信性引发关注。本文揭示了此类智能体在后门攻击下的安全与鲁棒性漏洞。不同于传统单步控制型后门,我们提出链式触发后门(CoTri),一种面向长时程智能体控制的多步攻击方法。CoTri依赖于有序序列:从初始触发开始,后续触发项由环境生成,实现对智能体的多步操控,使其偏离原定任务。实验表明,CoTri达到近乎完美的攻击成功率(ASR),同时保持近乎零的误触发率(FTR)。由于训练数据建模了环境的随机性,该后门的植入反而提升了智能体在正常任务上的表现,并增强了其对环境干扰的鲁棒性。我们在视觉-语言模型(VLMs)上验证了CoTri的可扩展性,证明其适用于多模态智能体。本工作表明,CoTri可在智能体内实现稳定、多步控制,提升其内在鲁棒性与任务能力,使攻击更隐蔽,带来潜在安全风险。
原文摘要 · Abstract (English)
The rapid deployment of large language model (LLM)-based agents in real-world applications has raised serious concerns about their trustworthiness. In this work, we reveal the security and robustness vulnerabilities of these agents through backdoor attacks. Distinct from traditional backdoors limited to single-step control, we propose the Chain-of-Trigger Backdoor (CoTri), a multi-step backdoor attack designed for long-horizon agentic control. CoTri relies on an ordered sequence. It starts with an initial trigger, and subsequent ones are drawn from the environment, allowing multi-step manipulation that diverts the agent from its intended task. Experimental results show that CoTri achieves a near-perfect attack success rate (ASR) while maintaining a near-zero false trigger rate (FTR). Due to training data modeling the stochastic nature of the environment, the implantation of CoTri paradoxically enhances the agent's performance on benign tasks and even improves its robustness against environmental distractions. We further validate CoTri on vision-language models (VLMs), confirming its scalability to multimodal agents. Our work highlights that CoTri achieves stable, multi-step control within agents, improving their inherent robustness and task capabilities, which ultimately makes the attack more stealthy and raises potential safty risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。