arXiv:2506.19389cs.CV2025-06被引 1

视觉语言模型读取图像文字能力突然涌现,早于语义理解但晚于图文匹配。

Emergence of Text Readability in Vision Language Models

  • 训练后期文本可读性突然出现,而语义理解逐步发展。
  • 图文匹配与文本识别相比更慢,需更强语义整合能力。
  • 提示需针对性训练策略,加速模型对文字的稳健理解。

我们研究了视觉语言模型(VLMs)在训练过程中识别图像内文本内容的能力如何涌现。分析揭示了一个关键现象:图像中的文本可读性在经过大量训练迭代后突然出现,而语义理解则从训练初期就开始渐进式发展。这种延迟涌现可能反映出对比学习最初优先关注通用语义理解,文本特定的符号处理随后才发展。有趣的是,将图像与渲染文本匹配的能力发展得更慢,表明需要更深层次的语义整合。这些发现凸显了设计定制化训练策略以加速VLM中稳健文本理解的必要性,为未来优化多模态学习的研究奠定了基础。

原文摘要 · Abstract (English)

We investigate how the ability to recognize textual content within images emerges during the training of Vision-Language Models (VLMs). Our analysis reveals a critical phenomenon: the ability to read textual information in a given image \textbf{(text readability)} emerges abruptly after substantial training iterations, in contrast to semantic content understanding which develops gradually from the early stages of training. This delayed emergence may reflect how contrastive learning tends to initially prioritize general semantic understanding, with text-specific symbolic processing developing later. Interestingly, the ability to match images with rendered text develops even slower, indicating a deeper need for semantic integration. These findings highlight the need for tailored training strategies to accelerate robust text comprehension in VLMs, laying the groundwork for future research on optimizing multimodal learning.

视觉语言模型文本识别多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。