用电子书作参考,大模型修正古籍错字,提升越南语古籍识别精度。
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
- 以电子书为参考,用大模型修正历史文档的错别字和漏字。
- 生成的伪标签数据集平均得分8.72,优于现有最佳模型的7.03。
- 适合古籍数字化、语言学研究及历史文本修复领域使用。
从年代久远的古籍中提取带变音符号的语言文本仍具挑战性,原因包括意外失真、时间导致的退化以及缺乏数据集。尽管已有独立拼写纠错方法,但在历史文献上表现有限,因错误组合多样且现代与古典语料分布差异大。本文提出利用可用的内容型电子书作为参考基底,结合大语言模型纠正不完整OCR输出。该方法生成高精度的伪页对页标签,尤其针对变音符号在历史条件下的识别难题。流程可消除多种噪声,解决缺字、漏词及顺序错乱问题。基于此构建的古典越南语书籍大规模OCR数据集,在10分制评测中平均得分8.72,显著优于当前最优的基于Transformer的越南语拼写纠错模型(7.03)。同时训练了基准OCR模型,实验结果表明其性能优于广泛使用的开源方案。相关数据集将公开发布,支持后续研究。
原文摘要 · Abstract (English)
Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While standalone spell correction approaches have been proposed, they show limited performance for historical documents due to numerous possible OCR error combinations and differences between modern and classical corpus distributions. We propose a method utilizing available content-focused ebooks as a reference base to correct imperfect OCR-generated text, supported by large language models. This technique generates high-precision pseudo-page-to-page labels for diacritic languages, where small strokes pose significant challenges in historical conditions. The pipeline eliminates various types of noise from aged documents and addresses issues such as missing characters, words, and disordered sequences. Our post-processing method, which generated a large OCR dataset of classical Vietnamese books, achieved a mean grading score of 8.72 on a 10-point scale. This outperformed the state-of-the-art transformer-based Vietnamese spell correction model, which scored 7.03 when evaluated on a sampled subset of the dataset. We also trained a baseline OCR model to assess and compare it with well-known engines. Experimental results demonstrate the strength of our baseline model compared to widely used open-source solutions. The resulting dataset will be released publicly to support future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。