用合成数据提升马拉地语文字识别,性能远超现有方法
LV-ROVER-MLT: Low-Resource Maltese OCR by Synthetic Fine-Tuning and Multi-Stream Arbitration

- 用合成数据微调Tesseract,融合五路识别流
- 比赛成绩CER仅0.0074,优于第二名近一半
- 适合低资源语言文字识别研究者参考
马耳他语拥有大量文本语料和预训练语言模型,但段落级文字识别训练数据稀缺;NOMOCRAT仅提供57页标注数据。LV-ROVER-MLT结合合成微调的Tesseract 5与五路互补识别流,采用适配马耳他语变音符号和连字符的词级仲裁机制。在DocEng 2026马耳他语文字识别竞赛中,该系统以0.0074的留出集词错误率(CER)排名第一,次优方案为0.0161,NOMOCRAT为0.0163。相同方法在卢森堡语上也显著优于原始Tesseract,匈牙利语结果不明确。研究构建了包含36,803对样本的马耳他语段落级OCR语料库,来源为EUR-Lex和Wikipedia。代码、模型权重及数据均已公开。
原文摘要 · Abstract (English)
Maltese has substantial text corpora and pretrained language models, but paragraph-scale OCR training data remains scarce; NOMOCRAT provides 57 verified annotated pages. LV-ROVER-MLT combines synthetic fine-tuning of Tesseract~5 with five complementary recognition streams and lexicon-gated word-level arbitration adapted to Maltese diacritics and hyphenation. In the DocEng~2026 Maltese OCR competition, the system placed first with held-out CER 0.0074; the next-ranked submission scored 0.0161 and NOMOCRAT scored 0.0163. The same approach produced a significant improvement over stock Tesseract on Luxembourgish, while the Hungarian result was inconclusive. A 36,803-pair Maltese OCR corpus constructed from EUR-Lex and Wikipedia provides an additional paragraph-level resource. Code, model weights, and corpus data are public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。