arXiv:2605.04808cs.AI2026-05被引 7

首个可控制的交互式红队平台,用于评估AI代理的安全性。

DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents

论文配图:DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
图 1 · 摘自论文原文
  • 构建跨14个真实场景的可控红队平台,模拟主流系统环境。
  • 提出自主红队代理DTap-Red,自动发现多种攻击路径并验证效果。
  • 生成大规模评测数据集,揭示代理系统性漏洞,适合安全研究者使用。

AI代理在复杂任务中日益广泛应用,但其高能力也带来显著安全风险。现实案例显示,攻击者可轻易诱导代理泄露密钥、删除数据或发起未经授权操作。由于代理运行于动态且不可信环境中,包含外部工具、异构数据源和频繁用户交互,其安全性评估极为困难。现有方法缺乏可控制、可复现的大规模风险评估环境。为此,我们提出首个可控制且交互式的红队平台DTap,覆盖14个真实领域和50余个仿真环境,复现Google Workspace、PayPal、Slack等常用系统。为实现规模化评估,我们进一步提出DTap-Red——首个自主红队代理,能系统探索提示注入、工具滥用、技能组合等多种攻击向量,并根据恶意目标自动生成有效攻击策略。基于DTap-Red,我们构建了大型红队评测数据集DTap-Bench,包含跨领域高质量样本及可验证的评判机制,用于自动验证攻击结果。通过DTap,我们对多个主流代理进行了大规模评估,涵盖不同基础模型、安全策略、风险类别与攻击方式,揭示系统性漏洞模式,为下一代安全代理研发提供关键洞见。

原文摘要 · Abstract (English)

AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their high capability and flexibility, such agents raise significant security and safety concerns. A growing number of real-world incidents have shown that adversaries can easily manipulate agents into performing harmful actions, such as leaking API keys, deleting user data, or initiating unauthorized transactions. Evaluating agent security is inherently challenging, as agents operate in dynamic, untrusted environments involving external tools, heterogeneous data sources, and frequent user interactions. However, realistic, controllable, and reproducible environments for large-scale risk assessment remain largely underexplored. To address this gap, we introduce the DecodingTrust-Agent Platform (DTap), the first controllable and interactive red-teaming platform for AI agents, spanning 14 real-world domains and over 50 simulation environments that replicate widely used systems such as Google Workspace, Paypal, and Slack. To scale the risk assessment of agents in DTap, we further propose DTap-Red, the first autonomous red-teaming agent that systematically explores diverse injection vectors (e.g., prompt, tool, skill, environment, combinations) and autonomously discovers effective attack strategies tailored to varying malicious goals. Using DTap-Red, we curate DTap-Bench, a large-scale red-teaming dataset comprising high-quality instances across domains, each paired with a verifiable judge to automatically validate attack outcomes. Through DTap, we conduct large-scale evaluations of popular AI agents built on various backbone models, spanning security policies, risk categories, and attack strategies, revealing systematic vulnerability patterns and providing valuable insights for developing secure next-generation agents.

AI安全红队测试智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。