视觉语言模型在字体识别上表现不佳,易受文字内容干扰。
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
- 构建字体识别基准FRB,含易难两版测试集
- 主流模型准确率低,且受文字语义干扰严重
- 注意力分析揭示模型难以捕捉字体语义特征
现代视觉语言模型(VLMs)在图像识别等任务中表现优异,但在细粒度识别任务中的能力仍存疑问。日常设计场景中,用户常需识别文本所用字体。尽管许多VLM具备多模态能力且免费可用,其是否真能识别字体仍不明。为此,我们提出字体识别基准FRB,包含15种常用字体,分两个版本:(i) 简单版,10个句子以不同字体渲染;(ii) 困难版,每个样本为15种字体名称本身,引入斯特鲁普效应,挑战模型感知能力。对多种VLM的评估显示:(i) 当前模型字体识别能力有限,多数先进模型性能不理想,且易受文字语义干扰;(ii) 少样本学习与思维链提示(CoT)对提升准确率帮助甚微;(iii) 注意力分析揭示模型难以捕捉字体的语义特征。
原文摘要 · Abstract (English)
Modern Vision-Language Models (VLMs) exhibit remarkable visual and linguistic capabilities, achieving impressive performance in various tasks such as image recognition and object localization. However, their effectiveness in fine-grained tasks remains an open question. In everyday scenarios, individuals encountering design materials, such as magazines, typography tutorials, research papers, or branding content, may wish to identify aesthetically pleasing fonts used in the text. Given their multimodal capabilities and free accessibility, many VLMs are often considered potential tools for font recognition. This raises a fundamental question: Do VLMs truly possess the capability to recognize fonts? To investigate this, we introduce the Font Recognition Benchmark (FRB), a compact and well-structured dataset comprising 15 commonly used fonts. FRB includes two versions: (i) an easy version, where 10 sentences are rendered in different fonts, and (ii) a hard version, where each text sample consists of the names of the 15 fonts themselves, introducing a stroop effect that challenges model perception. Through extensive evaluation of various VLMs on font recognition tasks, we arrive at the following key findings: (i) Current VLMs exhibit limited font recognition capabilities, with many state-of-the-art models failing to achieve satisfactory performance and being easily affected by the stroop effect introduced by textual information. (ii) Few-shot learning and Chain-of-Thought (CoT) prompting provide minimal benefits in improving font recognition accuracy across different VLMs. (iii) Attention analysis sheds light on the inherent limitations of VLMs in capturing semantic features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。