用表情包序列发动零扰动攻击,让NLP模型误判却不易被察觉。
Emoti-Attack: Zero-Perturbation Adversarial Attacks on NLP Systems via Emoji Sequences
- 利用表情符号作为独立攻击层,实现隐蔽扰动。
- 在大小模型上均有效,攻击成功率高且不明显改变语义。
- 适合研究模型安全、对抗样本防御的学者与工程师。
深度神经网络在自然语言处理领域取得显著成功,推动了如ChatGPT等广泛应用。然而,这些模型对对抗攻击的脆弱性仍是重大隐患。与图像等连续域不同,文本存在于离散空间,导致句、词、字符级别的微小修改易被人类感知,且文本不可微,使传统优化方法难以应用。以往文本对抗攻击研究集中于字符级、词级、句子级或多层级方法,但普遍存在效率低或语义破坏明显的缺陷。本文提出新型攻击方法Emoti-Attack,通过操控表情符号序列生成细微而有效的扰动。不同于传统字符或词级策略,该方法将表情符号视为独立攻击维度,扰动更隐蔽,对文本语义影响极小。该方向此前未受重视,多数研究仅将表情符号插入视为字符级攻击的延伸。实验表明,Emoti-Attack在大、小模型上均表现强劲,是提升NLP系统对抗鲁棒性的有力工具。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) have achieved remarkable success in the field of natural language processing (NLP), leading to widely recognized applications such as ChatGPT. However, the vulnerability of these models to adversarial attacks remains a significant concern. Unlike continuous domains like images, text exists in a discrete space, making even minor alterations at the sentence, word, or character level easily perceptible to humans. This inherent discreteness also complicates the use of conventional optimization techniques, as text is non-differentiable. Previous research on adversarial attacks in text has focused on character-level, word-level, sentence-level, and multi-level approaches, all of which suffer from inefficiency or perceptibility issues due to the need for multiple queries or significant semantic shifts. In this work, we introduce a novel adversarial attack method, Emoji-Attack, which leverages the manipulation of emojis to create subtle, yet effective, perturbations. Unlike character- and word-level strategies, Emoji-Attack targets emojis as a distinct layer of attack, resulting in less noticeable changes with minimal disruption to the text. This approach has been largely unexplored in previous research, which typically focuses on emoji insertion as an extension of character-level attacks. Our experiments demonstrate that Emoji-Attack achieves strong attack performance on both large and small models, making it a promising technique for enhancing adversarial robustness in NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。