arXiv:2508.14557cs.CVcs.LG2025-08被引 3

利用文档内字符冗余提升低质量印刷文本的识别准确率

Improving OCR using internal document redundancy

  • 通过字符形状冗余构建无监督纠错模型
  • 在退化文档上显著提升OCR识别率,尤其适用于历史档案
  • 适合处理老旧印刷品、档案扫描件等低质量文本

当前基于深度学习的OCR系统虽具备一定泛化能力,但在低质量印刷文档识别上仍表现不佳,尤其在跨域数据差异大而域内变化小的情况下。现有方法未能充分利用文档内部的字符形状冗余。本文提出一种无监督方法,通过引入改进的高斯混合模型(GMM),结合期望最大化(EM)算法与簇内重对齐过程及正态性统计检验,有效利用文档内字符重复模式来修正OCR输出并优化聚类。实验表明,该方法在多种退化程度的文档上均取得显著提升,包括恢复乌拉圭军方档案以及17世纪至20世纪中期的欧洲报纸文本。

原文摘要 · Abstract (English)

Current OCR systems are based on deep learning models trained on large amounts of data. Although they have shown some ability to generalize to unseen data, especially in detection tasks, they can struggle with recognizing low-quality data. This is particularly evident for printed documents, where intra-domain data variability is typically low, but inter-domain data variability is high. In that context, current OCR methods do not fully exploit each document's redundancy. We propose an unsupervised method by leveraging the redundancy of character shapes within a document to correct imperfect outputs of a given OCR system and suggest better clustering. To this aim, we introduce an extended Gaussian Mixture Model (GMM) by alternating an Expectation-Maximization (EM) algorithm with an intra-cluster realignment process and normality statistical testing. We demonstrate improvements in documents with various levels of degradation, including recovered Uruguayan military archives and 17th to mid-20th century European newspapers.

OCR文档修复无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。