arXiv:2606.09700cs.CRcs.HC2026-06中稿 · publication at USE…

用视觉技巧骗过AI,让有害内容逃过检测却仍被人类识别

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

论文配图:What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks
图 1 · 摘自论文原文
  • 通过调整字间距、加粗等排版方式,让文本对人眼显眼但对模型隐蔽
  • 仅用3次查询即可生成86%人类识别率、检测率低于1%的对抗样本
  • 揭示当前AI审核系统忽视人类视觉感知的深层缺陷,适合安全与伦理研究者

基于大语言模型(LLM)的内容审核系统是防范网络有害内容的关键防线。然而,这些系统主要依赖分词后的文本,往往忽略人类在理解内容时自然使用的视觉线索。我们发现这一局限导致根本性漏洞:人类能轻易识别为有害的内容可能完全避开自动化审核。为此,我们提出人类可感知对抗攻击(HPAA),通过具有视觉显著性的排版操作(如间距、强调、空间布局)将有害表达嵌入看似无害的文本中,既保持人类可读性又降低机器检出率。该攻击在黑盒环境下仅需少量查询(3次)即可自动生成规避内容,无需模型访问或梯度信息。我们在多个数据集和十三个主流审核系统(包括商业API与前沿开源防护墙)上验证,生成内容的人类识别率超86%,而各系统检测率均低于1%。我们进一步分析了促成成功规避的排版因素,揭示现有审核架构无法捕捉此类信号的原因,并探讨可行防御方案。研究结果揭示了当前基于LLM的审核系统在人类感知对齐上的根本盲点,呼吁更贴近人类认知的审核机制。

原文摘要 · Abstract (English)

Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally use when interpreting content. We show that this limitation creates a fundamental vulnerability: content readily recognized as harmful by humans can evade automated moderation. To systematically study this problem, we introduce Human-Perceptible Adversarial Attacks (HPAA), which embed harmful expressions into otherwise benign text using visually salient typographic manipulations. HPAA strategically combines features such as spacing, emphasis, and spatial arrangement to preserve human recognition while reducing machine detectability. Operating in a black-box setting with a small query budget, the attack automatically generates evasive content without model access or gradient information. We evaluate HPAA on multiple datasets and thirteen widely deployed moderation systems, including commercial APIs and state-of-the-art open-source guardrails. With only three detector queries, generated attacks achieve over 86\% human recognition while keeping detection rates below 1\% across evaluated systems. We further identify the typographic factors driving successful evasion, analyze why current moderation architectures fail to capture these signals, and discuss practical defenses. Our findings reveal a fundamental blind spot in current LLM-based moderation systems and motivate moderation approaches that better align with human perceptual understanding.

对抗攻击内容审核视觉感知LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。