用推理和符号编码绕过LLM安全机制,95%成功率攻破GPT系列模型
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- 通过推理解析中间步骤,让模型自行‘想出’有害行为
- 对GPT系列攻击成功率超95%,全目标平均达70%
- 揭示安全调优与有用性之间的根本矛盾,适合安全研究者参考
大型语言模型在多项任务中表现出色,但其被用于有害目的的风险仍令人担忧。为强化防御,需探究利用模型架构与学习范式固有缺陷的通用越狱攻击。本文提出新型越狱技术HaPLa,仅需黑盒访问目标模型。该方法包含两项核心策略:1)反事实推理框架,引导模型推断有害行为的合理中间步骤,而非直接响应显式有害请求;2)符号编码,一种轻量灵活的内容混淆方式,因当前模型仍主要对明确有害关键词敏感。实验显示,HaPLa在GPT系列模型上攻击成功率超过95%,所有目标平均达70%。进一步分析不同符号编码规则表明,不显著降低模型对良性查询的帮助性前提下,难以实现安全调优。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their potential misuse for harmful purposes remains a significant concern. To strengthen defenses against such vulnerabilities, it is essential to investigate universal jailbreak attacks that exploit intrinsic weaknesses in the architecture and learning paradigms of LLMs. In response, we propose \textbf{H}armful \textbf{P}rompt \textbf{La}undering (HaPLa), a novel and broadly applicable jailbreaking technique that requires only black-box access to target models. HaPLa incorporates two primary strategies: 1) \textit{abductive framing}, which instructs LLMs to infer plausible intermediate steps toward harmful activities, rather than directly responding to explicit harmful queries; and 2) \textit{symbolic encoding}, a lightweight and flexible approach designed to obfuscate harmful content, given that current LLMs remain sensitive primarily to explicit harmful keywords. Experimental results show that HaPLa achieves over 95% attack success rate on GPT-series models and 70% across all targets. Further analysis with diverse symbolic encoding rules also reveals a fundamental challenge: it remains difficult to safely tune LLMs without significantly diminishing their helpfulness in responding to benign queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。