用数学问题伪装恶意指令,突破大模型安全防护。
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

- 将有害内容转化为形式化数学问题绕过安全检测
- 攻击成功率高达46%~56%,跨模型通用
- 新模型更抗攻击,但仍有漏洞,适合安全研究者
大型语言模型虽有安全机制防止有害输出,但主要依赖语义模式匹配。本文发现,将有害提示编码为连贯的数学问题(如集合论、形式逻辑、量子力学),可在八种目标模型和两个基准上实现46%至56%的平均攻击成功率。关键在于,攻击效果不取决于数学符号本身,而在于辅助大模型是否深度重构有害内容为真实数学问题;仅添加数学格式而不重构的规则编码,效果与未编码基线无异。本文提出新型形式逻辑编码,攻击成功率与集合论相当,证明该漏洞在多种数学形式下均存在。重复后处理实验表明,此类攻击对简单提示增强具有鲁棒性。值得注意的是,新模型(GPT-5、GPT-5-Mini)比旧模型更具鲁棒性,但仍可被攻破。研究揭示当前安全框架的根本缺陷,呼吁发展基于数学结构推理而非表面语义的防御机制。
原文摘要 · Abstract (English)
Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We show that encoding harmful prompts as coherent mathematical problems -- using formalisms such as set theory, formal logic, and quantum mechanics -- bypasses these filters at high rates, achieving 46%--56% average attack success across eight target models and two established benchmarks. Crucially, the effectiveness depends not on mathematical notation itself, but on whether a helper LLM deeply reformulates the harmful content into a genuine mathematical problem: rule-based encodings that apply mathematical formatting without such reformulation perform no better than unencoded baselines. We introduce a novel Formal Logic encoding that achieves attack success comparable to Set Theory, demonstrating that this vulnerability generalizes across mathematical formalisms. Additional experiments with repeat post-processing confirm that these attacks are robust to simple prompt augmentation. Notably, newer models (GPT-5, GPT-5-Mini) show substantially greater robustness than older models, though they remain vulnerable. Our findings highlight fundamental gaps in current safety frameworks and motivate defenses that reason about mathematical structure rather than surface-level semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。