arXiv:2502.01205cs.CL2025-02被引 18

用大模型修正古籍文字识别错误,效果因语言而异。

OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

  • 用开源大模型对历史文献的识别错误进行后处理
  • 英文字符错误率下降,芬兰语效果不理想
  • 揭示大模型在古籍文本纠错中的实际局限性

光学字符识别(OCR)系统在转录历史文献时常引入错误,为提升文本质量,需进行后处理校正。本研究评估了开源大模型在英语和芬兰语历史文献上的OCR错误纠正效果。我们探索了参数优化、量化、片段长度影响及文本续写方法等多种策略。结果表明,尽管现代大模型在英文上显著降低了字符错误率(CER),但在芬兰语上未能达到实用水平。研究揭示了大模型在大规模历史语料库中进行OCR后处理时的潜力与局限。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study evaluates the use of open-weight LLMs for OCR error correction in historical English and Finnish datasets. We explore various strategies, including parameter optimization, quantization, segment length effects, and text continuation methods. Our results demonstrate that while modern LLMs show promise in reducing character error rates (CER) in English, a practically useful performance for Finnish was not reached. Our findings highlight the potential and limitations of LLMs in scaling OCR post-correction for large historical corpora.

OCR纠错大模型历史文献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。