arXiv:2409.17699cs.CRcs.AI2024-09AAAI被引 9

用简单统计方法构建专家混合模型,高效识别提示攻击

MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks

  • 采用语言统计特征组合多个专家模型进行检测
  • 在不误伤正常请求的前提下识别90%的攻击
  • 计算开销极低,适合实际部署于大模型系统

大型语言模型在各类应用中的普及凸显了防范潜在越狱攻击的紧迫性。此类攻击利用模型漏洞,威胁数据完整性和用户隐私。防护机制(Guardrails)是关键防御手段,但现有方案在检测准确率与计算效率方面仍存不足。本文强调大模型越狱防护的重要性,突出输入层防护机制的价值。提出一种新型防护架构MoJE(Mixture of Jailbreak Experts),通过简单的语言统计技术,在推理阶段保持极低计算开销的同时,显著提升越狱攻击检测能力。实验表明,该方法可有效检测90%的越狱攻击,且不影响正常提示的处理,显著增强大模型对越狱攻击的防御能力。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) in diverse applications underscores the pressing need for robust security measures to thwart potential jailbreak attacks. These attacks exploit vulnerabilities within LLMs, endanger data integrity and user privacy. Guardrails serve as crucial protective mechanisms against such threats, but existing models often fall short in terms of both detection accuracy, and computational efficiency. This paper advocates for the significance of jailbreak attack prevention on LLMs, and emphasises the role of input guardrails in safeguarding these models. We introduce MoJE (Mixture of Jailbreak Expert), a novel guardrail architecture designed to surpass current limitations in existing state-of-the-art guardrails. By employing simple linguistic statistical techniques, MoJE excels in detecting jailbreak attacks while maintaining minimal computational overhead during model inference. Through rigorous experimentation, MoJE demonstrates superior performance capable of detecting 90% of the attacks without compromising benign prompts, enhancing LLMs security against jailbreak attacks.

大模型安全越狱攻击输入防护轻量检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。