arXiv:2606.29649cs.CL2026-06

高分辨率让AI难识别恶意ASCII艺术,中文英文都一样。

Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages

论文配图:Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages
图 1 · 摘自论文原文
  • 测试8种字符构造模式在10个分辨率下的检测效果
  • 超过特定分辨率后,检测率大幅下降,最高达70%失效
  • 基于文字的构造方式最难被发现,适合做绕过攻击

大型视觉语言模型(VLMs)被广泛用于内容审核,但容易受到“越狱攻击”——即通过将有害文本以ASCII艺术形式呈现来规避检测。本文研究图像分辨率对VLM在八种字符构造模式(L1-L8,从密集块字符到嵌入文字设计)下检测有害ASCII艺术的影响。在英、中语料库上,采用生成十种分辨率图像的流水线,评估八种先进VLM的表现,探究是否存在跨模型、模式与语言的统一检测失败阈值。结果显示,当分辨率超过某阈值后,检测率急剧下降;且基于文字的构造模式在整个分辨率范围内最难以被检测。该发现揭示了VLM内容审核系统的系统性漏洞,呼吁建立考虑分辨率的评估标准。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) are increasingly deployed as content moderation tools, yet they remain vulnerable to jailbreak attacks in which harmful text is visually encoded as ASCII art. This can allow inappropriate or harmful content to bypass moderation systems. To address this vulnerability, this paper investigates how image resolution affects VLM detection of harmful ASCII art across eight character construction modes (L1-L8), ranging from dense block characters to word-embedded designs. We evaluate eight state-of-the-art VLMs on English and Chinese corpora using a pipeline that generates ASCII art images at ten resolution scales, probing whether a consistent detection-failure threshold exists across models, modes, and languages. Results indicate that detection rates decline sharply above certain resolution thresholds, and that word-based modes are the most resistant to detection across the full resolution range. These findings reveal a systematic vulnerability in VLM-based content moderation systems and motivate resolution-aware evaluation standards.

内容审核ASCII艺术视觉语言模型安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。