arXiv:2608.11878cs.CRcs.CL2026-08

构建可扩展的对抗环境,测试大模型工具代理的安全性。

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

论文配图:ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
图 1 · 摘自论文原文
  • 用模拟器自动生成带漏洞的交互环境
  • 发现攻击时机与位置影响成功率
  • 适合安全研究者和模型对齐团队使用

集成外部工具的大语言模型(LLM)代理易受嵌入环境状态中的间接提示注入攻击。现有研究多依赖人工构建或复用环境、随机的LLM工具模拟及预设注入点,难以在更广泛领域实现可扩展的安全研究。为此,我们提出ToolHazard——一种可扩展的对抗环境合成框架,减少人工干预并支持新增种子领域与算力扩展。该框架包含环境模拟器、攻击者代理和用户模拟器,能生成可执行的状态化环境,发现可行的注入点并生成环境特定的攻击载荷,构建基于状态的长期任务。基于此,我们构建了ToolHazard-Bench,用于在复杂工作流和多样化环境攻击下压力测试代理。实验揭示了代理显著漏洞,并表明注入时机与位置影响攻击效果。此外,由ToolHazard生成的对齐数据在ToolHazard-Bench和AgentDojo上均提升了安全性,同时保持良性任务性能。

原文摘要 · Abstract (English)

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.

大模型安全对抗攻击智能体自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。