用字符画绕过大模型安全机制,一次攻击即可成功
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- 通过预测试确定字符画识别最优参数,实现精准攻击
- 在4个主流开源模型上成功触发越狱,攻击成功率超90%
- 可迁移至GPT-4o等商用模型,对内容过滤器具强鲁棒性
大型语言模型(LLMs)在计算机应用中的融合带来了变革性能力,但也引发重大安全挑战。现有安全对齐主要依赖语义理解,使模型易受非标准数据表示的攻击。本文提出ArtPerception,一种新型黑盒越狱框架,利用ASCII艺术绕过当前最先进的(SOTA)LLMs的安全措施。不同于依赖迭代暴力攻击的旧方法,ArtPerception采用系统化的两阶段策略:第一阶段进行一次性、模型特定的预测试,实证确定最佳的ASCII艺术识别参数;第二阶段利用这些洞察,发起高效的一次性恶意越狱攻击。我们提出改进的莱文斯坦距离(MLD)度量,以更精细评估模型的识别能力。在四个SOTA开源模型上的全面实验中,验证了其优越的越狱性能。进一步通过实验证明该框架在真实场景下的适用性,成功迁移至GPT-4o、Claude Sonnet 3.7和DeepSeek-V3等主流商业模型,并对LLaMA Guard与Azure内容过滤器进行了严格有效性分析。研究结果表明,真正的模型安全需防御文本输入中多模态解释空间,凸显战略性侦察型攻击的有效性。内容警告:本文包含潜在有害及冒犯性模型输出。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into computer applications has introduced transformative capabilities but also significant security challenges. Existing safety alignments, which primarily focus on semantic interpretation, leave LLMs vulnerable to attacks that use non-standard data representations. This paper introduces ArtPerception, a novel black-box jailbreak framework that strategically leverages ASCII art to bypass the security measures of state-of-the-art (SOTA) LLMs. Unlike prior methods that rely on iterative, brute-force attacks, ArtPerception introduces a systematic, two-phase methodology. Phase 1 conducts a one-time, model-specific pre-test to empirically determine the optimal parameters for ASCII art recognition. Phase 2 leverages these insights to launch a highly efficient, one-shot malicious jailbreak attack. We propose a Modified Levenshtein Distance (MLD) metric for a more nuanced evaluation of an LLM's recognition capability. Through comprehensive experiments on four SOTA open-source LLMs, we demonstrate superior jailbreak performance. We further validate our framework's real-world relevance by showing its successful transferability to leading commercial models, including GPT-4o, Claude Sonnet 3.7, and DeepSeek-V3, and by conducting a rigorous effectiveness analysis against potential defenses such as LLaMA Guard and Azure's content filters. Our findings underscore that true LLM security requires defending against a multi-modal space of interpretations, even within text-only inputs, and highlight the effectiveness of strategic, reconnaissance-based attacks. Content Warning: This paper includes potentially harmful and offensive model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。