提出视觉文本纠错任务,揭示大模型在图像文字理解上的致命短板
ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

- 设计真实图像文本纠错任务,检验模型对图文上下文的理解能力
- 顶尖模型与人类仍存显著差距,多数模型误判文字内容
- 为开发更懂文字的视觉语言模型提供新基准
视觉语言模型(VLMs)在通用视觉任务中表现优异,但在理解图像中的文字方面仍存在明显不足。本文提出ReViCo(Real Visual Correction),一个通过视觉文本错误纠正任务评估VLM文本理解能力的新基准。该任务要求模型识别并修复真实图像中的文字错误,需深入理解文字与其周围视觉上下文的互动关系。我们采用提示引导策略和针对性模型训练两种范式对多个VLM进行评测,结果揭示即使最先进的模型与人类相比仍有显著性能差距,且多数模型难以准确感知视觉文本,导致频繁纠错错误。通过揭示这些局限,ReViCo为构建更鲁棒、更注重文本的VLMs提供了新的基准基础。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。