构建可执行红队测试框架,精准评估大模型代理安全漏洞
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

- 基于安全约束生成攻击,通过沙箱执行并验证危害结果
- 六模型三框架下平均攻击成功率65.69%,证据可见性影响评估结果
- 发现'认知-执行鸿沟',无需训练的提醒策略可降低70%以上违规
大型语言模型代理通过结合语言推理与外部工具执行复杂任务。对抗输入可能利用代理与其环境的交互,导致执行中违反安全策略。现有评估常将代理安全简化为单一攻击成功率(ASR),忽略了暴露、执行、观察与裁决环节,可能混淆实际违规与证据可见性。我们提出REDAgentBench,一个可执行的自主红队测试与精准测量框架。该框架从明确的安全约束及代理系统漏洞生成攻击,在隔离服务沙箱中运行,并通过服务响应与最终状态变化验证危害效果。基准包含跨五个服务面的1,661个案例。在六种模型和三种代理调度器上,宏观平均ASR为65.69%;报告的ASR随调度器与证据视角变化,评估上下文披露会改变执行行为。在状态基础诊断组中,近五分之一经确认的违规发生在代理声明相关约束或风险之后,揭示了‘认知-执行鸿沟’。最后,一种无训练的策略提醒在匹配重播中使确认违规减少超过70个百分点。这些发现表明,可执行评估能提升安全测量精度,并识别可操作的干预点。
原文摘要 · Abstract (English)
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAgentBench, an executable framework for autonomous red-teaming and faithful measurement. It derives attacks from explicit safety constraints and associated agent-system vulnerabilities, runs them in isolated service sandboxes, and verifies harmful effects from service receipts and final-state changes. The benchmark contains 1,661 cases across five service surfaces. Across six models and three agent harnesses, macro-average ASR is 65.69%; reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior. In a state-grounded diagnostic cohort, almost one in five confirmed violations with resolved action anchors occurs after the agent states the relevant constraint or risk, revealing a Recognition--Execution Gap. Finally, a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay. These findings show that executable evaluation can improve safety measurement and identify actionable intervention points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。