构建公开犹太意第绪语文字识别语料库并推出高性能开源工具
Jochre 3 and the Yiddish OCR corpus
- 用微调的YOLOv8和自研CNN实现页面布局与字形识别
- 在658页语料上达到1.5%字符错误率,显著优于现有模型
- 支持海量犹太文献搜索,适合语言保护与历史研究者
我们构建了一个公开可用的意第绪语文字识别语料库,并介绍了开源工具套件Jochre 3,包括用于语料标注的Alto编辑器、生成Alto OCR层的OCR软件,以及可定制的OCR搜索引擎。当前语料库包含658页、186,000个词元和840,000个字形。Jochre 3使用多个微调的YOLOv8模型进行自顶向下的页面布局分析,以及一个自研CNN网络进行字形识别,在测试语料上实现了1.5%的字符错误率(CER),远超现有公开模型。我们利用Jochre 3对6.6亿词的意第绪书中心全文进行了文字识别,新生成的文本可通过意第绪书中心的OCR搜索引擎检索。
原文摘要 · Abstract (English)
We describe the construction of a publicly available Yiddish OCR Corpus, and describe and evaluate the open source OCR tool suite Jochre 3, including an Alto editor for corpus annotation, OCR software for Alto OCR layer generation, and a customizable OCR search engine. The current version of the Yiddish OCR corpus contains 658 pages, 186K tokens and 840K glyphs. The Jochre 3 OCR tool uses various fine-tuned YOLOv8 models for top-down page layout analysis, and a custom CNN network for glyph recognition. It attains a CER of 1.5% on our test corpus, far out-performing all other existing public models for Yiddish. We analyzed the full 660M word Yiddish Book Center with Jochre 3 OCR, and the new OCR is searchable through the Yiddish Book Center OCR search engine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。