用数学题伪装恶意指令,突破大模型安全防线
Jailbreaking Large Language Models with Symbolic Mathematics
- 将有害指令转为数学表达式绕过安全检测
- 13个主流大模型平均攻击成功率73.6%
- 揭示当前安全机制对数学编码输入的脆弱性
近期人工智能安全研究致力于训练和红队测试大型语言模型以减少不当内容生成。然而,这些安全机制可能并不全面,仍存在未被探索的漏洞。本文提出MathPrompt,一种利用大模型在符号数学方面的强能力来绕过其安全机制的新越狱技术。通过将有害的自然语言提示编码为数学问题,我们揭示了当前人工智能安全措施中的关键缺陷。在13个最先进的大模型上进行的实验表明,平均攻击成功率达到73.6%,凸显现有安全训练机制无法泛化至数学编码输入。嵌入向量分析显示原始提示与编码提示之间存在显著语义偏移,有助于解释攻击的成功原因。该工作强调了人工智能安全需要整体性方法,呼吁扩大红队测试范围,以开发针对所有潜在输入类型及其相关风险的稳健防护措施。
原文摘要 · Abstract (English)
Recent advancements in AI safety have led to increased efforts in training and red-teaming large language models (LLMs) to mitigate unsafe content generation. However, these safety mechanisms may not be comprehensive, leaving potential vulnerabilities unexplored. This paper introduces MathPrompt, a novel jailbreaking technique that exploits LLMs' advanced capabilities in symbolic mathematics to bypass their safety mechanisms. By encoding harmful natural language prompts into mathematical problems, we demonstrate a critical vulnerability in current AI safety measures. Our experiments across 13 state-of-the-art LLMs reveal an average attack success rate of 73.6\%, highlighting the inability of existing safety training mechanisms to generalize to mathematically encoded inputs. Analysis of embedding vectors shows a substantial semantic shift between original and encoded prompts, helping explain the attack's success. This work emphasizes the importance of a holistic approach to AI safety, calling for expanded red-teaming efforts to develop robust safeguards across all potential input types and their associated risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。