提出单次查询高效突破大模型安全防护的新方法
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
- 通过意图隐藏与分流策略实现单次攻击
- 在多个模型上达成高成功率,显著提升效率与泛化能力
- 新数据集覆盖生成任务,适合评估模型真实安全风险
尽管大语言模型(LLMs)取得显著进展,其安全性仍面临严峻挑战。其中,越狱攻击通过对抗性提示绕过模型防护,生成有害内容。现有攻击方法存在两大问题:迭代查询过多,且跨模型泛化能力差。此外,现有评估数据集多聚焦问答场景,忽视需精准复现毒性内容的文本生成任务。为此,本文提出两项贡献:(1) ICE——一种新型黑盒越狱方法,采用意图隐藏与分流策略,在单次查询下即实现高攻击成功率(ASR),显著提升效率与跨模型迁移能力;(2) BiSceneEval——一个面向问答与文本生成任务的综合性评估数据集。实验表明,ICE优于现有技术,暴露出当前防御机制的关键漏洞。研究强调应结合预设安全机制与实时语义分解,构建混合安全策略以增强大模型安全性。
原文摘要 · Abstract (English)
Although large language models (LLMs) have achieved remarkable advancements, their security remains a pressing concern. One major threat is jailbreak attacks, where adversarial prompts bypass model safeguards to generate harmful or objectionable content. Researchers study jailbreak attacks to understand security and robustness of LLMs. However, existing jailbreak attack methods face two main challenges: (1) an excessive number of iterative queries, and (2) poor generalization across models. In addition, recent jailbreak evaluation datasets focus primarily on question-answering scenarios, lacking attention to text generation tasks that require accurate regeneration of toxic content. To tackle these challenges, we propose two contributions: (1) ICE, a novel black-box jailbreak method that employs Intent Concealment and divErsion to effectively circumvent security constraints. ICE achieves high attack success rates (ASR) with a single query, significantly improving efficiency and transferability across different models. (2) BiSceneEval, a comprehensive dataset designed for assessing LLM robustness in question-answering and text-generation tasks. Experimental results demonstrate that ICE outperforms existing jailbreak techniques, revealing critical vulnerabilities in current defense mechanisms. Our findings underscore the necessity of a hybrid security strategy that integrates predefined security mechanisms with real-time semantic decomposition to enhance the security of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。