开源19世纪希腊文文献语料库,突破古希腊文光学识别难题
The Patrologia Graeca Corpus: OCR, Annotation, and Open Release of Noisy Nineteenth-Century Polytonic Greek Editions
- 用YOLO+CRNN管道处理复杂排版与退化字体
- 字符错误率仅1.05%,词错误率4.69%,性能领先
- 适合古典学、语言学与AI训练,含标注与布局信息
我们发布首个大规模开放的19世纪古希腊文语料库——Patrologia Graeca Corpus,涵盖未数字化的《帕特罗洛吉亚希腊语集》(PG)残余卷册。这些文本采用复杂的双语(希腊语-拉丁语)排版,且希腊文使用高度退化的多音调符号系统。通过结合基于YOLO的版面检测与基于CRNN的文本识别的专用流程,我们实现1.05%的字符错误率(CER)和4.69%的词错误率(WER),显著优于现有古希腊文光学识别系统。该语料库包含约六百万个经过词元化与词性标注的词项,并配有完整的OCR与版面标注。除学术价值外,本资源为噪声多音调希腊文的OCR设立了新基准,并可作为未来大模型训练的数据基础。
原文摘要 · Abstract (English)
We present the Patrologia Graeca Corpus, the first large-scale open OCR and linguistic resource for nineteenthcentury editions of Ancient Greek. The collection covers the remaining undigitized volumes of the Patrologia Graeca (PG), printed in complex bilingual (Greek-Latin) layouts and characterized by highly degraded polytonic Greek typography. Through a dedicated pipeline combining YOLO-based layout detection and CRNN-based text recognition, we achieve a character error rate (CER) of 1.05% and a word error rate (WER) of 4.69%, largely outperforming existing OCR systems for polytonic Greek. The resulting corpus contains around six million lemmatized and part-of-speech tagged tokens, aligned with full OCR and layout annotations. Beyond its philological value, this corpus establishes a new benchmark for OCR on noisy polytonic Greek and provides training material for future models, including LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。