测试视觉化文字对视觉语言模型理解能力的影响
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
- 对比纯文本与图像中的视觉化文本,评估模型表现差异
- 30多个模型显示:视觉化文本使性能显著下降
- 适合研究多模态理解、模型鲁棒性的学者参考
视觉语言模型(VLMs)在跨模态理解任务中表现优异,但现有基准大多聚焦于纯文本查询。在真实场景中,语言常以图像中的视觉化文本形式出现,这引发疑问:当前VLMs能否同等处理此类输入?我们提出VISTA-Bench,一个涵盖多模态感知、推理到单模态理解的系统性基准。通过在受控渲染条件下对比纯文本与视觉化文本问题,评估模型对视觉化文本的理解能力。对超过30个代表性VLM的广泛评估显示,尽管语义一致,模型在视觉化文本上的表现显著下降,且感知难度越高,差距越明显。该结果揭示了模型对渲染变化的敏感性,凸显其在文本与像素间统一表征的不足。VISTA-Bench为诊断此局限提供可复现框架,推动更鲁棒的跨模态理解发展。数据集与代码已公开于https://github.com/QingAnLiu/VISTA-Bench。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios, language also frequently appears as visualized text embedded in images, raising the question of whether current VLMs handle such input requests comparably. We introduce VISTA-Bench, a systematic benchmark from multimodal perception, reasoning, to unimodal understanding domains. It evaluates visualized text understanding by contrasting pure-text and visualized-text questions under controlled rendering conditions. Extensive evaluation of over 30 representative VLMs reveals a pronounced modality gap: models that perform well on pure-text queries often degrade substantially when equivalent semantic content is presented as visualized text. This gap is further amplified by increased perceptual difficulty, highlighting sensitivity to rendering variations despite unchanged semantics. Overall, VISTA-Bench provides a principled evaluation framework to diagnose this limitation and to guide progress toward more unified language representations across tokenized text and pixels. The source dataset and code are publicly available at https://github.com/QingAnLiu/VISTA-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。