arXiv:2602.14524cs.CV2026-02被引 4

对比两种模型在古籍识别中的错误模式,揭示其对学术研究的潜在影响。

Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model

  • 用长度加权准确率和错误分析法,比较TrOCR与Qwen在历史文本上的表现。
  • Qwen虽整体错误率低但会隐性修正古体字形,可能改变历史语义。
  • TrOCR更保真但易产生连锁错误,适合关注文字原貌的研究者。

十八世纪印刷文本的光学字符识别(OCR)因印刷质量退化、古体字形及非标准化拼写而极具挑战。尽管基于Transformer的OCR系统和视觉语言模型(VLMs)在整体准确率上表现良好,但字符错误率(CER)和词错误率(WER)难以反映其在学术使用中的可靠性。本文通过长度加权准确率和假设驱动的错误分析,比较了专用OCR Transformer(TrOCR)与通用视觉语言模型(Qwen)在行级历史英文文本上的表现。结果表明,Qwen虽具备更低的CER/WER且对退化输入更鲁棒,却表现出选择性语言规整与拼写标准化,可能无声地改变具有历史意义的文本形式;而TrOCR更一致地保持拼写真实性,但更容易发生错误级联传播。研究发现,模型架构的归纳偏置会系统性地塑造错误结构。即使整体准确率相近,不同模型在错误局部性、可检测性和下游学术风险方面仍存在显著差异,强调在历史数字化流程中需进行架构感知的评估。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) of eighteenth-century printed texts remains challenging due to degraded print quality, archaic glyphs, and non-standardized orthography. Although transformer-based OCR systems and Vision-Language Models (VLMs) achieve strong aggregate accuracy, metrics such as Character Error Rate (CER) and Word Error Rate (WER) provide limited insight into their reliability for scholarly use. We compare a dedicated OCR transformer (TrOCR) and a general-purpose Vision-Language Model (Qwen) on line-level historical English texts using length-weighted accuracy metrics and hypothesis driven error analysis. While Qwen achieves lower CER/WER and greater robustness to degraded input, it exhibits selective linguistic regularization and orthographic normalization that may silently alter historically meaningful forms. TrOCR preserves orthographic fidelity more consistently but is more prone to cascading error propagation. Our findings show that architectural inductive biases shape OCR error structure in systematic ways. Models with similar aggregate accuracy can differ substantially in error locality, detectability, and downstream scholarly risk, underscoring the need for architecture-aware evaluation in historical digitization workflows.

OCR历史文本视觉语言模型错误分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。