用特殊字体伪装恶意文本,让模型误判却仍能被人读懂。
Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
- 利用字体差异制造人类可读、模型误判的对抗样本
- 轻量版高效,强力版攻击效果显著,跨模型通用
- 威胁文本生成、语音合成等多模态任务,警示安全风险
随着社交媒体发展,用户使用风格化字体和类字体表情符号表达个性,生成视觉吸引但人类可读的文本。然而,这些字体在NLP模型中引入隐藏漏洞:人类易读,模型却将字符视为不同标记,造成干扰。我们识别出这一人机感知差异,提出基于风格的攻击方法Style Attack Disguise(SAD)。设计两种版本:轻量版提升查询效率,强力版实现更优攻击性能。在情感分类与机器翻译任务上,针对传统模型、大语言模型及商用服务的实验表明SAD具有强攻击效果。同时验证其对文本到图像、文本到语音生成等多模态任务的潜在威胁。
原文摘要 · Abstract (English)
With social media growth, users employ stylistic fonts and font-like emoji to express individuality, creating visually appealing text that remains human-readable. However, these fonts introduce hidden vulnerabilities in NLP models: while humans easily read stylistic text, models process these characters as distinct tokens, causing interference. We identify this human-model perception gap and propose a style-based attack, Style Attack Disguise (SAD). We design two sizes: light for query efficiency and strong for superior attack performance. Experiments on sentiment classification and machine translation across traditional models, LLMs, and commercial services demonstrate SAD's strong attack performance. We also show SAD's potential threats to multimodal tasks including text-to-image and text-to-speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。