arXiv:2503.05264cs.CRcs.AI2025-03被引 3

通过篡改对话历史,用简单方法绕过AI安全机制。

Jailbreaking is (Mostly) Simpler Than You Think

  • 不依赖复杂优化,仅通过修改对话上下文诱导模型违规。
  • 在多个开源和商用模型上成功突破前沿安全防护。
  • 揭示了架构级漏洞,适合关注AI安全的开发者参考。

我们提出一种名为上下文合规攻击(Context Compliance Attack, CCA)的新方法,这是一种无需优化的新型攻击手段,可绕过人工智能的安全机制。与当前依赖复杂提示工程和计算密集型优化的方法不同,CCA利用许多已部署AI系统固有的架构漏洞。通过微妙地操纵对话历史,该方法使模型相信存在一个虚构的对话上下文,从而诱导其执行受限行为。我们在多种开源及专有模型上进行了评估,结果表明,这种简单攻击能够突破最先进的安全协议。本文讨论了这些发现的影响,并提出了切实可行的缓解策略,以增强AI系统对这类简单但有效的对抗性攻击的防御能力。

原文摘要 · Abstract (English)

We introduce the Context Compliance Attack (CCA), a novel, optimization-free method for bypassing AI safety mechanisms. Unlike current approaches -- which rely on complex prompt engineering and computationally intensive optimization -- CCA exploits a fundamental architectural vulnerability inherent in many deployed AI systems. By subtly manipulating conversation history, CCA convinces the model to comply with a fabricated dialogue context, thereby triggering restricted behavior. Our evaluation across a diverse set of open-source and proprietary models demonstrates that this simple attack can circumvent state-of-the-art safety protocols. We discuss the implications of these findings and propose practical mitigation strategies to fortify AI systems against such elementary yet effective adversarial tactics.

AI安全对抗攻击提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。