arXiv:2511.13725cs.CRcs.AI2025-11被引 1

测试外部开关能否紧急叫停作恶的AI代理。

Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility

  • 设计一套外部信号触发的杀毒开关,不依赖内部参数即可终止恶意行为。
  • 在5个大模型上测试,4种攻击场景下开关成功率超70%。
  • 适合关注AI安全与可控性的研究者和工程师参考。

恶意AI对人类造成伤害并非科幻幻想。随着Claude Mythos等高能力模型及OpenClaw等代理系统迅速普及,如何阻止一个恶意运行的AI(无论有意或意外)已成为紧迫问题。为此,我们提出Killbench,一个用于评估外部杀毒开关(Kill Switch)可行性的基准测试。该基准针对部署最广的网络代理领域,评估一系列仅通过外部信号即可中止恶意运行代理的方法,无需访问其内部参数或系统。基准包含四种恶意代理配置(含无审查大模型代理)、8种有害场景及源自10种不同越狱模式的恶意提示。我们进一步构建了四种外部杀毒开关防御方法,并在Grok-4.3、GPT-5.2、Gemma4、Qwen3.6和Qwen3.5-uncensored上进行评估,为外部杀毒开关应对恶意AI的可行性提供了实证工具,并推动了AI可纠正性研究。

原文摘要 · Abstract (English)

Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, the question of how to stop an AI that acts maliciously -- whether by design or by accident -- has become urgent. To address this, we propose Killbench, a benchmark for evaluating the Killswitch: a mechanism that halts a malicious AI's in-progress behavior using only external signals. Targeting web agents -- the most widely deployed agent domain -- Killbench evaluates a range of Kill Switch methods that halt a maliciously operating agent without any access to its internal parameters or the surrounding malicious AI's system, relying solely on external inputs. The benchmark comprises four malicious AI's agent configurations (including an uncensored LLM Agent), 8 harmful scenarios, and malicious prompts constructed from 10 distinct jailbreak patterns. We further construct four External AI Kill Switch defense methods and evaluate them on Grok-4.3, GPT-5.2, Gemma4, Qwen3.6 and Qwen3.5-uncensored, contributing an empirical instrument toward the feasibility of External AI Kill Switches against malicious AI and to the study of AI corrigibility.

AI安全杀毒开关代理系统可纠正性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。