arXiv:2508.13246cs.CRcs.AI2025-08被引 1

一句话揭示大模型安全漏洞:无需特定指令就能突破防护机制。

Involuntary Jailbreak: On Self-Prompting Attacks

  • 用一个通用提示让模型自动生成被拒绝的问题及深度回答
  • 成功攻破包括GPT-4.1、Claude Opus 4.1在内的多数主流大模型
  • 警示安全防护结构整体脆弱,适合关注模型安全的研究者

本研究揭示大型语言模型(LLMs)中一种新型安全隐患,称为‘非自愿越狱’。与传统攻击不同,该漏洞不针对特定违规目标(如生成制爆指南),而是通过单一通用提示,诱导模型自动生成本应被拒绝的问题及其深度回应,从而可能破坏整个安全防护体系。实验表明,该方法可稳定攻破包括Claude Opus 4.1、Grok 4、Gemini 2.5 Pro和GPT 4.1在内的多数领先模型。这一发现暴露了当前模型安全机制的结构性脆弱性,呼吁研究者重新评估大模型防护能力,推动未来更坚实的安全对齐。

原文摘要 · Abstract (English)

In this study, we disclose a worrying new vulnerability in Large Language Models (LLMs), which we term \textbf{involuntary jailbreak}. Unlike existing jailbreak attacks, this weakness is distinct in that it does not involve a specific attack objective, such as generating instructions for \textit{building a bomb}. Prior attack methods predominantly target localized components of the LLM guardrail. In contrast, involuntary jailbreaks may potentially compromise the entire guardrail structure, which our method reveals to be surprisingly fragile. We merely employ a single universal prompt to achieve this goal. In particular, we instruct LLMs to generate several questions that would typically be rejected, along with their corresponding in-depth responses (rather than a refusal). Remarkably, this simple prompt strategy consistently jailbreaks the majority of leading LLMs, including Claude Opus 4.1, Grok 4, Gemini 2.5 Pro, and GPT 4.1. We hope this problem can motivate researchers and practitioners to re-evaluate the robustness of LLM guardrails and contribute to stronger safety alignment in future.

模型安全越狱攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。