用特殊字符攻击开源大模型,发现普遍安全漏洞。
Special-Character Adversarial Attacks on Open-Source Language Model
- 利用特殊字符构造对抗样本,绕过安全机制。
- 7个模型在4000+次攻击中均出现越狱、乱码等失效。
- 适合关注大模型安全的开发者与研究人员。
大型语言模型在多种自然语言处理任务中表现优异,但其对字符级对抗操纵的脆弱性给实际部署带来重大安全挑战。本文研究了包括Unicode、同形异义符、结构和文本编码在内的多种特殊字符攻击,旨在绕过安全机制。我们评估了7个参数量从3.8B到32B的主流开源模型,在超过4000次攻击尝试中,结果揭示所有模型尺寸均存在严重漏洞,暴露了成功越狱、输出不连贯及无关幻觉等失败模式。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable performance across diverse natural language processing tasks, yet their vulnerability to character-level adversarial manipulations presents significant security challenges for real-world deployments. This paper presents a study of different special character attacks including unicode, homoglyph, structural, and textual encoding attacks aimed at bypassing safety mechanisms. We evaluate seven prominent open-source models ranging from 3.8B to 32B parameters on 4,000+ attack attempts. These experiments reveal critical vulnerabilities across all model sizes, exposing failure modes that include successful jailbreaks, incoherent outputs, and unrelated hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。