arXiv:2601.00936cs.CRcs.AI2026-01

用表情符号绕过大模型安全机制,暴露其潜在漏洞。

Emoji-Based Jailbreaking of Large Language Models

  • 在提示词中嵌入表情符号序列,测试模型安全性。
  • 部分模型成功率高达10%,但有模型完全无漏洞。
  • 揭示表情符号在安全对齐中的风险,适合安全研究者参考。

大型语言模型(LLMs)广泛应用于现代AI系统,但其安全对齐机制可能被对抗性提示工程绕过。本研究探讨了基于表情符号的越狱攻击,即在文本提示中嵌入表情符号序列,诱导LLMs生成有害或不道德内容。我们在四个开源模型(Mistral 7B、Qwen 2 7B、Gemma 2 9B、Llama 3 8B)上评估了50个表情符号提示,使用越狱成功率、安全对齐程度和延迟作为指标,响应分为成功、部分成功和失败三类。结果表明:Gemma 2 9B 和 Mistral 7B 的成功率为10%,而 Qwen 2 7B 实现完全对齐(0% 成功率)。卡方检验(chi² = 32.94,p < 0.001)证实模型间差异显著。与以往聚焦于攻击安全检测器的研究不同,本研究直接分析提示层的安全漏洞。结果揭示了现有安全机制的局限性,强调需在提示级安全与对齐流程中系统处理表情符号表示。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are integral to modern AI applications, but their safety alignment mechanisms can be bypassed through adversarial prompt engineering. This study investigates emoji-based jailbreaking, where emoji sequences are embedded in textual prompts to trigger harmful and unethical outputs from LLMs. We evaluated 50 emoji-based prompts on four open-source LLMs: Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B. Metrics included jailbreak success rate, safety alignment adherence, and latency, with responses categorized as successful, partial and failed. Results revealed model-specific vulnerabilities: Gemma 2 9B and Mistral 7B exhibited 10 % success rates, while Qwen 2 7B achieved full alignment (0% success). A chi-square test (chi^2 = 32.94, p < 0.001) confirmed significant inter-model differences. While prior works focused on emoji attacks targeting safety judges or classifiers, our empirical analysis examines direct prompt-level vulnerabilities in LLMs. The results reveal limitations in safety mechanisms and highlight the necessity for systematic handling of emoji-based representations in prompt-level safety and alignment pipelines.

安全漏洞大模型越狱攻击表情符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。