测试视觉语言模型对碎片化文字的识别能力,发现其远不如人类稳定。
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
- 用拼接重叠字符构造人类可读但模型难辨的文本
- 模型在扰动下准确率大幅下降,输出常无意义
- 适合关注多模态鲁棒性与跨文字系统应用的研究者
书写是普遍的文化技术,利用视觉进行符号化交流。人类具有显著韧性:即使字符断裂、粘连或部分遮挡,仍能轻松识别文字。本文探究先进视觉语言模型(VLMs)是否具备类似韧性。我们构建了两个受心理学实验启发的基准测试,覆盖中文表意文字和英文字母文字,通过拼接、重组和叠加字形生成对人类可读但对模型不可读的刺激。尽管模型在干净文本上表现良好,但在这些扰动下性能严重下降,常输出无关或不连贯内容。结果表明模型存在结构性缺陷:过度依赖通用视觉不变性,而忽视符号分割、组合与绑定所需的组成先验。我们公开了刺激生成代码、提示和评估协议,以支持透明复现与后续研究。研究呼吁设计更注重符号结构建模的架构与训练策略,并为教育、无障碍、文化遗产和安全等领域的多模态系统部署提供具体挑战。
原文摘要 · Abstract (English)
Writing is a universal cultural technology that reuses vision for symbolic communication. Humans display striking resilience: we readily recognize words even when characters are fragmented, fused, or partially occluded. This paper investigates whether advanced vision language models (VLMs) share this resilience. We construct two psychophysics inspired benchmarks across distinct writing systems, Chinese logographs and English alphabetic words, by splicing, recombining, and overlaying glyphs to yield ''visible but unreadable'' stimuli for models while remaining legible to humans. Despite strong performance on clean text, contemporary VLMs show a severe drop under these perturbations, frequently producing unrelated or incoherent outputs. The pattern suggests a structural limitation: models heavily leverage generic visual invariances but under rely on compositional priors needed for robust literacy. We release stimuli generation code, prompts, and evaluation protocols to facilitate transparent replication and follow up work. Our findings motivate architectures and training strategies that encode symbol segmentation, composition, and binding across scripts, and they delineate concrete challenges for deploying multimodal systems in education, accessibility, cultural heritage, and security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。