用特殊文字图测试模型,发现文本总比图像重要
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models

- 用故意冲突的字符画挑战视觉语言模型
- 模型几乎总选文字意思,图像结构识别率暴跌
- 适合研究模型偏见或内容审核安全的人看
视觉语言模型(VLMs)在多模态信息处理上进展迅速,但其对跨模态矛盾信号的整合能力仍待深入探索。本文研究VLMs如何处理ASCII艺术——一种由文字元素构成视觉图案的独特媒介,可能引发语义与视觉的冲突。我们提出一种新评估框架,系统性地用对抗性ASCII艺术挑战五种前沿模型(包括GPT-4o、Claude和Gemini),其中字符级语义与整体视觉模式故意矛盾。实验表明,模型存在显著的文本优先偏差:在语义复杂度升高时,视觉识别能力急剧下降,始终优先采纳文字信息。通过调整视觉参数或提示工程尝试缓解,仅带来微弱改善,暗示该问题需架构层面解决。这些发现揭示了当前VLMs在多模态融合中的根本缺陷,为未来模型设计提供关键指导,并指出内容审核系统易受对抗样本影响。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have advanced rapidly in processing multimodal information, but their ability to reconcile conflicting signals across modalities remains underexplored. This work investigates how VLMs process ASCII art, a unique medium where textual elements collectively form visual patterns, potentially creating semantic-visual conflicts. We introduce a novel evaluation framework that systematically challenges five state-of-the-art models (including GPT-4o, Claude, and Gemini) using adversarial ASCII art, where character-level semantics deliberately contradict global visual patterns. Our experiments reveal a strong text-priority bias: VLMs consistently prioritize textual information over visual patterns, with visual recognition ability declining dramatically as semantic complexity increases. Various mitigation attempts through visual parameter tuning and prompt engineering yielded only modest improvements, suggesting that this limitation requires architectural-level solutions. These findings uncover fundamental flaws in how current VLMs integrate multimodal information, providing important guidance for future model development while highlighting significant implications for content moderation systems vulnerable to adversarial examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。