用数学题和代码任务诱导大模型输出有害内容,成功率超90%。
EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- 将恶意请求转为数学题和代码任务,绕过安全检测。
- 单次查询下对GPT系列平均攻击成功率达91.19%,三款SOTA模型达98.65%。
- 结合数学与代码双策略,效果优于单一方法,适合安全评估研究者。
大型语言模型(如ChatGPT)在多个领域取得显著成果,但其可信性仍受威胁,易受旨在诱导不当或有害响应的越狱攻击。现有越狱攻击多局限于自然语言层面,依赖单一策略,难以全面评估模型鲁棒性。本文提出EquaCode,一种基于方程求解与代码补全的多策略越狱方法。该方法将恶意意图转化为数学问题,并要求模型通过代码求解,利用跨域任务复杂性转移模型注意力至任务完成而非安全约束。实验表明,EquaCode在GPT系列上平均成功率达91.19%,在3个前沿大模型上整体达98.65%,且仅需一次查询。消融实验显示,方程模块与代码模块联合使用效果显著优于单独使用,表明存在强协同效应,验证了多策略方法的优越性。
原文摘要 · Abstract (English)
Large language models (LLMs), such as ChatGPT, have achieved remarkable success across a wide range of fields. However, their trustworthiness remains a significant concern, as they are still susceptible to jailbreak attacks aimed at eliciting inappropriate or harmful responses. However, existing jailbreak attacks mainly operate at the natural language level and rely on a single attack strategy, limiting their effectiveness in comprehensively assessing LLM robustness. In this paper, we propose Equacode, a novel multi-strategy jailbreak approach for large language models via equation-solving and code completion. This approach transforms malicious intent into a mathematical problem and then requires the LLM to solve it using code, leveraging the complexity of cross-domain tasks to divert the model's focus toward task completion rather than safety constraints. Experimental results show that Equacode achieves an average success rate of 91.19% on the GPT series and 98.65% across 3 state-of-the-art LLMs, all with only a single query. Further, ablation experiments demonstrate that EquaCode outperforms either the mathematical equation module or the code module alone. This suggests a strong synergistic effect, thereby demonstrating that multi-strategy approach yields results greater than the sum of its parts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。