用表情符号测试大模型安全,发现文本评估可能遗漏真实漏洞
Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation
- 用表情符号增强提示词测试模型安全性
- 不同模型对表情符号攻击的抵抗能力差异显著
- 研究提醒评估需覆盖多种输入形式,适合安全研究者参考
大型语言模型(LLMs)的安全评估主要依赖文本对抗性提示,可能忽视其他输入形式带来的潜在漏洞。本文以表情符号增强的提示为测试案例,评估了4个开源模型(Mistral 7B、Qwen 2 7B、Gemma 2 9B、Llama 3 8B)在50个提示下的表现。结果显示模型鲁棒性差异明显:Gemma 2 9B 和 Mistral 7B 的成功攻击率分别为10%,Llama 3 8B 为6%,而 Qwen 2 7B 完全抵抗(0%)。卡方检验(χ² = 32.94, p < 0.001)证实各模型结果分布存在显著差异。研究表明,模型安全性对输入表示敏感,仅使用标准文本提示的评估可能低估实际漏洞。
原文摘要 · Abstract (English)
Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented prompts as a test case for this gap, evaluating 50 prompts across four open-source LLMs (Mistral 7B, Qwen 2 7B, Gemma 2 9B, Llama 3 8B). Results show substantial variation in robustness: Gemma 2 9B and Mistral 7B exhibit non-zero success rates (10%), Llama 3 8B 6%, while Qwen 2 7B shows complete resistance (0% success rate). A chi-square test ($χ^2 = 32.94, p < 0.001$) confirms significant differences in outcome distributions. These findings indicate that robustness is sensitive to input representation, and that evaluations restricted to standard text prompts may underrepresent model vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。