构建动态红队测试框架,统一评估各类自主智能系统的安全漏洞。
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

- 基于分层表示法,自动发现系统结构并执行自适应攻击
- 使用105种攻击探针覆盖多种攻击目标,验证45个异构系统
- 支持安全策略效果评估,适合安全研究人员和开发者使用
由大语言模型驱动的自主智能系统正快速发展为自主决策系统,暴露出传统大模型漏洞之外的新攻击路径。现有安全评估往往局限于特定实现或领域,难以跨系统统一比较。为此,我们提出RIFT-Bench,一种基于表征的动态红队测试方法,可对多样化自主架构进行统一评估。该方法包含两个自动化阶段:发现阶段提取系统结构,扫描阶段执行自适应对抗攻击。直接使用105个自适应攻击探针,覆盖多种攻击向量与目标,对45个不同实现的自主系统进行了验证,表明该方法能有效泛化至异构架构。此外,RIFT-Bench还可直接评估缓解策略的有效性。这些能力使其成为实践中可扩展的自主智能系统安全评估基础。基础设施代码与基准数据集见https://tinyurl.com/RIFTBench。
原文摘要 · Abstract (English)
Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems. To address this gap, we introduce RIFT-Bench, a representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures. Building on a novel hierarchical representation, RIFT-Bench operates in two automated phases: Discovery, which extracts system structure, and Scanning, which executes adaptive adversarial attacks. It directly evaluates the examined system using 105 adaptive adversarial probes spanning diverse attack vectors and objectives. We demonstrate the effectiveness of the proposed evaluation pipeline across 45 agentic systems spanning a diverse range of implementations, showing that the approach generalizes effectively to heterogeneous agentic architectures. Beyond systems and attacks, RIFT-Bench also supports direct evaluation of mitigation strategies. These key capabilities make RIFT-Bench a scalable foundation for security evaluation of agentic AI systems in practice. Infrastructure code and benchmark artifacts are available at https://tinyurl.com/RIFTBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。