用文字艺术绕过毒性检测,揭示现有系统重大漏洞
Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems
- 通过ASCII艺术构造视觉化文本,欺骗仅依赖文本的检测模型
- 在多个主流大模型和审核工具上实现100%攻击成功率
- 适合研究内容安全、对抗样本或模型鲁棒性的读者
我们提出一类新型对抗攻击,利用语言模型无法理解以ASCII艺术形式呈现的空间结构文本,从而规避毒性检测。为评估此类攻击效果,我们构建了ToxASCII基准,用于测试毒性检测系统对视觉混淆输入的鲁棒性。实验表明,该攻击在多种先进大语言模型及专用审核工具上均实现了100%的攻击成功率达到(ASR),暴露出当前纯文本审核系统存在显著漏洞。
原文摘要 · Abstract (English)
We introduce a novel class of adversarial attacks on toxicity detection models that exploit language models' failure to interpret spatially structured text in the form of ASCII art. To evaluate the effectiveness of these attacks, we propose ToxASCII, a benchmark designed to assess the robustness of toxicity detection systems against visually obfuscated inputs. Our attacks achieve a perfect Attack Success Rate (ASR) across a diverse set of state-of-the-art large language models and dedicated moderation tools, revealing a significant vulnerability in current text-only moderation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。