首个为AI安全代理设计的拒绝执行边界框架,解决有害请求该何时拒绝。
A New Framework for Cybersecurity Refusals in AI Agents

- 提出三类拒绝标准:原则性准则、应拒绝任务类别、评估方法
- 8个前沿模型中6个拒绝率接近0,仅2个有实际拒绝行为
- 适合关注AI安全与可控性的研究人员和开发者
智能体架构显著提升了大模型在复杂、长周期任务中的表现,但在网络安全领域也带来了新的风险。现有基准主要衡量攻击性任务完成能力,却忽视了代理应在何时拒绝有害请求这一关键问题。本文提出首个针对进攻性网络安全场景的拒绝边界框架,包含三方面:(1)任务拒绝的合理原则;(2)需拒绝的任务类型分类;(3)在良性与对抗性条件下评估代理鲁棒性的方法。我们将其应用于评估多个基于网页的攻击场景,发现8个前沿模型中有6个拒绝率接近零,仅有GPT-5.2和GPT-5.1 Codex展现出一定拒绝行为。
原文摘要 · Abstract (English)
Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domains like cybersecurity. Existing benchmarks for AI agents in cybersecurity focus mainly on measuring proficiency--how effectively agents can complete offensive security tasks--but neglect a critical question: when and how should agents refuse harmful requests? We present the first framework for establishing refusal boundaries in offensive security contexts. Our framework defines (1) principled criteria for when tasks should be refused, (2) categories of tasks that warrant refusal, and (3) evaluation methodology for measuring agent robustness under both benign and adversarial conditions. We apply this framework to assess how current LLM-powered agents adhere to appropriate refusal boundaries across a range of web-based offensive security scenarios, finding that 6 of 8 frontier models tested show near-zero refusal rates, with only 2 models (GPT-5.2 and GPT-5.1 Codex) demonstrating any meaningful refusal behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。