arXiv:2602.16901cs.AI2026-02被引 23

首个评估大模型代理长期攻击风险的基准测试

AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks

  • 设计五类新型长周期攻击,覆盖28个真实场景
  • 644个测试用例显示现有代理仍极易受攻击
  • 揭示单轮防御策略对长期攻击无效,适合安全研究者

大语言模型代理被越来越多地部署于长期、复杂的环境中解决难题,但这种扩展使其暴露于利用多轮用户-代理-环境交互实现单轮不可行目标的长期攻击中。为衡量代理对此类风险的脆弱性,我们提出AgentLAB,这是首个专门评估大语言模型代理对自适应长期攻击敏感性的基准。当前AgentLAB支持五种新型攻击类型:意图劫持、工具链攻击、任务注入、目标漂移和记忆污染,涵盖28个现实的代理环境和644个安全测试用例。基于AgentLAB,我们评估了代表性大语言模型代理,发现它们对长期攻击仍高度脆弱;此外,针对单轮交互设计的防御措施无法可靠缓解长期威胁。我们预期AgentLAB将在实际场景中追踪大语言模型代理安全进展方面发挥重要作用。该基准已公开发布于https://tanqiujiang.github.io/AgentLAB_main。

原文摘要 · Abstract (English)

LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user-agent-environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The benchmark is publicly available at https://tanqiujiang.github.io/AgentLAB_main.

大模型安全长周期攻击代理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。