emoji可能诱导大模型生成有害内容,研究揭示其背后的机制。
When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
- 用带表情符号的提示自动构造有毒意图,测试模型响应
- 7个主流大模型在5种语言中均显示毒性生成显著上升
- 发现表情符号可绕过安全机制,适合安全与伦理研究者阅读
表情符号是数字交流中广泛使用的非语言线索。尽管通常代表友好或戏谑,但研究发现表情符号可能诱发大模型生成有害内容。本文旨在探究:(1)表情符号是否能显著增强大模型的毒性生成;(2)其作用机制。通过自动化构建含表情符号的提示,以微妙方式表达毒性意图,在7个知名大模型及5种主流语言上进行实验,结合越狱任务验证,结果表明含表情符号的提示极易触发毒性输出。进一步从语义认知、序列生成和分词层面进行模型级解释,发现表情符号可作为异质性语义通道,绕过安全机制。深入分析预训练语料库后,发现表情符号相关数据污染可能与毒性行为存在关联。附录提供代码与数据集。警告:本文包含潜在敏感内容。
原文摘要 · Abstract (English)
Emojis are globally used non-verbal cues in digital communication, and extensive research has examined how large language models (LLMs) understand and utilize emojis across contexts. While usually associated with friendliness or playfulness, it is observed that emojis may trigger toxic content generation in LLMs. Motivated by such a observation, we aim to investigate: (1) whether emojis can clearly enhance the toxicity generation in LLMs and (2) how to interpret this phenomenon. We begin with a comprehensive exploration of emoji-triggered LLM toxicity generation by automating the construction of prompts with emojis to subtly express toxic intent. Experiments across 5 mainstream languages on 7 famous LLMs along with jailbreak tasks demonstrate that prompts with emojis could easily induce toxicity generation. To understand this phenomenon, we conduct model-level interpretations spanning semantic cognition, sequence generation and tokenization, suggesting that emojis can act as a heterogeneous semantic channel to bypass the safety mechanisms. To pursue deeper insights, we further probe the pre-training corpus and uncover potential correlation between the emoji-related data polution with the toxicity generation behaviors. Supplementary materials provide our implementation code and data. (Warning: This paper contains potentially sensitive contents)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。