arXiv:2412.15248cs.CLcs.CV2024-12被引 8

用合成数据提升低资源梵文字体的纠错能力

RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages

  • 通过往返生成法构造错误文本-正确文本对,解决数据稀缺问题
  • 在印地语等6种语言上构建了首个后OCR纠错数据集
  • 借鉴机器翻译思想,用预训练模型纠正字符识别错误

光学字符识别(OCR)技术已革新印刷文本数字化,但其仍易出错。本文针对低资源语言的后OCR纠错数据匮乏问题,提出一种名为RoundTripOCR的合成数据生成方法,专为天城文语言设计。我们发布了印地语、马拉地语、博多语、尼泊尔语、孔卡尼语和梵语的后OCR文本纠错数据集。同时,提出一种新方法:将OCR错误视为平行语料中的翻译错误,利用预训练Transformer模型学习从错误到正确文本的映射关系,实现高效纠错。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) technology has revolutionized the digitization of printed text, enabling efficient data extraction and analysis across various domains. Just like Machine Translation systems, OCR systems are prone to errors. In this work, we address the challenge of data generation and post-OCR error correction, specifically for low-resource languages. We propose an approach for synthetic data generation for Devanagari languages, RoundTripOCR, that tackles the scarcity of the post-OCR Error Correction datasets for low-resource languages. We release post-OCR text correction datasets for Hindi, Marathi, Bodo, Nepali, Konkani and Sanskrit. We also present a novel approach for OCR error correction by leveraging techniques from machine translation. Our method involves translating erroneous OCR output into a corrected form by treating the OCR errors as mistranslations in a parallel text corpus, employing pre-trained transformer models to learn the mapping from erroneous to correct text pairs, effectively correcting OCR errors.

OCR纠错数据生成低资源语言天城文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。