arXiv:2501.11623cs.CVcs.AI2025-01被引 28

大模型比传统OCR更准,能更好还原古籍手写文本。

Early evidence of how LLMs outperform traditional systems on OCR/HTR tasks for historical records

  • 用分行或整图输入,让大模型逐行识别手写表格数据。
  • 两轮提示下GPT-4o在分行输入中错误率最低,Claude Sonnet 3.5在整图输入表现最佳。
  • 结果接近人工标注,适合历史文献数字化项目使用。

我们评估了GPT-4o和Claude Sonnet 3.5两款大模型在将历史手写文档转录为表格格式方面的表现,并与EasyOCR、Keras、Pytesseract和TrOCR等传统OCR/HTR系统进行对比。针对表格数据的特性,分别设计了按行分割图像与使用完整扫描图作为输入的两种实验。基于字符错误率(CER)和BLEU分数,结果显示大模型性能优于传统方法。此外,通过与人工评估结果对比,进一步分析了整体扫描实验中影响CER和BLEU的关键因素。综合各项评价指标,最终得出结论:对分行图像采用两轮提示的GPT-4o,以及对整图输入采用两轮提示的Claude Sonnet 3.5,生成的转录结果最接近真实文本。

原文摘要 · Abstract (English)

We explore the ability of two LLMs -- GPT-4o and Claude Sonnet 3.5 -- to transcribe historical handwritten documents in a tabular format and compare their performance to traditional OCR/HTR systems: EasyOCR, Keras, Pytesseract, and TrOCR. Considering the tabular form of the data, two types of experiments are executed: one where the images are split line by line and the other where the entire scan is used as input. Based on CER and BLEU, we demonstrate that LLMs outperform the conventional OCR/HTR methods. Moreover, we also compare the evaluated CER and BLEU scores to human evaluations to better judge the outputs of whole-scan experiments and understand influential factors for CER and BLEU. Combining judgments from all the evaluation metrics, we conclude that two-shot GPT-4o for line-by-line images and two-shot Claude Sonnet 3.5 for whole-scan images yield the transcriptions of the historical records most similar to the ground truth.

大模型古籍识别文本转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。