用简单替换骗过仇恨言论检测模型,86.8%成功率
All You Need is "Leet": Evading Hate-speech Detection AI
- 用字符替换生成扰动,黑盒攻击主流检测模型
- 86.8%的仇恨文本可绕过检测,原意几乎不变
- 适合研究对抗攻击或内容安全的开发者
社交媒体和在线论坛日益普及,但也被用于传播仇恨言论。本文设计了黑盒攻击技术,通过生成扰动来欺骗基于深度学习的仇恨言论检测模型,降低其检测效率,同时尽量保持原始仇恨文本语义不变。最佳扰动攻击在86.8%的仇恨文本上成功绕过了检测系统。
原文摘要 · Abstract (English)
Social media and online forums are increasingly becoming popular. Unfortunately, these platforms are being used for spreading hate speech. In this paper, we design black-box techniques to protect users from hate-speech on online platforms by generating perturbations that can fool state of the art deep learning based hate speech detection models thereby decreasing their efficiency. We also ensure a minimal change in the original meaning of hate-speech. Our best perturbation attack is successfully able to evade hate-speech detection for 86.8 % of hateful text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。