通过图像修复与语义纠错结合,显著提升古籍文档的识别准确率。
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
- 先合成带退化的文档图像训练修复模型,再用ByT5修正识别错误
- 在13,831页历史文档上,字符错误率降低63.9%至70.3%
- 适合古籍数字化、档案整理等需要高精度文本提取的场景
本文提出PreP-OCR,一个两阶段流程:首先基于纯文本合成带多种字体、版式和退化操作的文档图像对,利用多方向块提取融合技术训练图像修复模型;其次使用在合成历史文本对上微调的ByT5模型,进行后处理纠错。在包含13,831页英文、法文、西班牙文历史文档的真实数据集上,该流程相较原始图像直接识别,字符错误率降低63.9%至70.3%。结果表明,图像修复与语言级纠错结合可有效提升古籍数字化中的文本提取质量。
原文摘要 · Abstract (English)
This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded historical documents. First, we synthesize document-image pairs from plaintext, rendering them with diverse fonts and layouts and then applying a randomly ordered set of degradation operations. An image restoration model is trained on this synthetic data, using multi-directional patch extraction and fusion to process large images. Second, a ByT5 post-OCR model, fine-tuned on synthetic historical text pairs, addresses remaining OCR errors. Detailed experiments on 13,831 pages of real historical documents in English, French, and Spanish show that the PreP-OCR pipeline reduces character error rates by 63.9-70.3% compared to OCR on raw images. Our pipeline demonstrates the potential of integrating image restoration with linguistic error correction for digitizing historical archives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。