arXiv:2607.24077cs.CVcs.LG2026-07

VLM OCR虽字符错误率低,却常生成语义失真文本,影响历史档案可信度。

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

论文配图:When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents
图 1 · 摘自论文原文
  • 用视觉语言模型做历史文档识别,比传统OCR更准
  • 即使字符错误率很低,仍会误改人名地名等关键信息
  • 适合关注真实语义准确性的历史文献数字化研究者

光学字符识别(OCR)是历史档案数字化的关键。近年来,视觉语言模型(VLMs)在标准基准上表现优异,成为传统OCR的有力替代。本文在来自乌拉圭独裁时期微缩胶片扫描的Berrutti数据集上,对比传统OCR与基于VLM的方法。尽管VLM在字符错误率(CER)和词错误率(WER)上持续领先,但定性分析揭示了标准指标无法捕捉的系统性缺陷:拼写规范化、虚构内容生成及语义替换(虽流畅但含义改变)。命名实体错误尤为严重,可能造成重大语义扭曲而对CER/WER影响甚微。该研究揭示了量化性能与实际转录保真度之间的关键差距,强调需建立超越字符级准确性的评估框架,以衡量转录的语义可靠性。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.

OCR视觉语言模型历史档案语义错误

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。