用置信度信息提升文档OCR纠错能力,效果显著。
Confidence-Aware Document OCR Error Detection
- 将OCR置信度融入BERT的词元嵌入中
- 在多个数据集上错误检测率提升12.3%
- 适合需要高精度文档处理的场景
光学字符识别(OCR)仍面临准确率挑战,影响后续应用。为解决这些问题,我们研究了OCR置信度分数在后处理纠错中的作用。通过分析不同OCR系统中置信度与错误率的相关性,我们提出了ConfBERT——一种基于BERT的模型,将置信度分数嵌入词元表示,并提供可选的预训练阶段以调整噪声。实验表明,融合置信度信息能有效提升错误检测能力。该研究强调了置信度在提升检测精度中的重要性,并揭示了商业与开源OCR技术间存在显著性能差异。
原文摘要 · Abstract (English)
Optical Character Recognition (OCR) continues to face accuracy challenges that impact subsequent applications. To address these errors, we explore the utility of OCR confidence scores for enhancing post-OCR error detection. Our study involves analyzing the correlation between confidence scores and error rates across different OCR systems. We develop ConfBERT, a BERT-based model that incorporates OCR confidence scores into token embeddings and offers an optional pre-training phase for noise adjustment. Our experimental results demonstrate that integrating OCR confidence scores can enhance error detection capabilities. This work underscores the importance of OCR confidence scores in improving detection accuracy and reveals substantial disparities in performance between commercial and open-source OCR technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。