构建日文票据OCR纠错基准,提升票据识别准确率
JaPOC: Japanese Post-OCR Correction Benchmark using Vouchers

- 基于真实票据数据构建日文OCR纠错基准数据集
- 使用语言模型基线显著提升整体识别准确率
- 为日本票据自动化处理提供可复用的评估标准
本文针对日文票据在光学字符识别(OCR)系统中的错误纠正问题,构建了首个公开可用的后OCR纠错基准。由于印章等噪声干扰,票据文本的准确识别面临挑战,而现有研究缺乏专门针对日文票据的纠错评估基准。研究首先测量了现有OCR服务在日文票据上的识别准确率,随后建立了包含真实场景票据的数据集,并提出基于语言模型的简单纠错基线方法。实验表明,所提方法能有效纠正识别错误,显著提升整体识别精度。
原文摘要 · Abstract (English)
In this paper, we create benchmarks and assess the effectiveness of error correction methods for Japanese vouchers in OCR (Optical Character Recognition) systems. It is essential for automation processing to correctly recognize scanned voucher text, such as the company name on invoices. However, perfect recognition is complex due to the noise, such as stamps. Therefore, it is crucial to correctly rectify erroneous OCR results. However, no publicly available OCR error correction benchmarks for Japanese exist, and methods have not been adequately researched. In this study, we measured text recognition accuracy by existing services on Japanese vouchers and developed a post-OCR correction benchmark. Then, we proposed simple baselines for error correction using language models and verified whether the proposed method could effectively correct these errors. In the experiments, the proposed error correction algorithm significantly improved overall recognition accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。