VLM在古希腊文献OCR中常凭语言习惯猜字,而非看图识别。
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions

- 用图像扰动和分词级分析,检验模型是否真看图
- 通用VLM错得像真的一样,但专精模型已不依赖图像
- 纠错需事后处理,实时干预无效,适合可解释性研究者
针对低资源古希腊校勘文本的光学字符识别(OCR),对比开放权重的视觉语言模型(VLMs)与传统OCR基线发现,VLM生成的错误常具语言流畅性,而传统引擎仅产生局部噪声。为分析解码过程中的视觉证据依赖,引入可控图像扰动及基于条件与无图像解码分布的分词级定位度量。字符级扰动下,VLM输出显著偏离真实文本,而传统OCR仍保持相对忠实;但分词级分析显示,模型依赖性存在差异:专用OCR模型在生成流利错误时几乎不依赖图像,而通用VLM即便出错仍受视觉输入影响。解码时干预无法可靠恢复视觉对齐,而事后语言模型修正仅能修复生成后文本。结果将OCR中语言先验依赖现象扩展至低资源历史文档与更广泛模型,表明流畅输出未必视觉对齐,呼吁以可解释性评估超越整体准确率。
原文摘要 · Abstract (English)
Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on language priors. Comparing open-weight VLMs with traditional OCR baselines on low-resource Ancient Greek critical editions, we show that VLM errors often remain fluent even when wrong, producing plausible Greek substitutions where traditional engines produce local recognition noise. To analyze visual evidence during decoding, we introduce controlled image perturbations and token-level grounding measures based on conditional versus image-free decoding distributions. Under character-level perturbations, VLMs diverge sharply from the perturbed ground truth while traditional OCR remains comparatively faithful; however, token-level analysis shows that prior reliance is model-specific: in an OCR-specialist model, fluent lexical errors are produced with little reliance on the image, whereas general-purpose VLMs remain conditioned on the visual input even when wrong. Decode-time interventions fail to reliably restore grounding, while post-OCR language-model correction improves several systems only by repairing text after generation. Our results extend prior evidence of OCR language-prior reliance to low-resource historical documents and a broader set of models, showing that fluent output is not necessarily visually grounded and motivating interpretability-driven evaluation beyond aggregate accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。