arXiv:2501.07300cs.CLcs.CV2025-01被引 1

提升萨米语文档识别准确率,为北欧原住民文献数字化提供可靠工具。

Comparative analysis of optical character recognition methods for Sámi texts from the National Library of Norway

  • 对比微调Transkribus、Tesseract和TrOCR三种主流OCR模型
  • 微调后Transkribus与TrOCR在萨米语数据上准确率显著优于Tesseract
  • 结合人工标注、机器标注与合成图像,小样本下也能实现高精度识别

光学字符识别(OCR)对挪威国家图书馆(NLN)的数字化进程至关重要,可将扫描文档转为可机器处理的文本。然而,现有系统对萨米语文档的识别准确率不足。由于OCR质量直接影响后续应用,有必要评估并改进针对萨米语言的识别方法。本文对三种成熟OCR方法——Transkribus、Tesseract和TrOCR进行微调与评估。结果表明,在NLN萨米语数据集上,Transkribus与TrOCR表现优于Tesseract;而Tesseract在域外数据集上表现更优。此外,通过微调预训练模型,并结合人工标注、机器标注及合成文本图像,即使仅有少量人工标注数据,也可实现高精度萨米语识别。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) is crucial to the National Library of Norway's (NLN) digitisation process as it converts scanned documents into machine-readable text. However, for the Sámi documents in NLN's collection, the OCR accuracy is insufficient. Given that OCR quality affects downstream processes, evaluating and improving OCR for text written in Sámi languages is necessary to make these resources accessible. To address this need, this work fine-tunes and evaluates three established OCR approaches, Transkribus, Tesseract and TrOCR, for transcribing Sámi texts from NLN's collection. Our results show that Transkribus and TrOCR outperform Tesseract on this task, while Tesseract achieves superior performance on an out-of-domain dataset. Furthermore, we show that fine-tuning pre-trained models and supplementing manual annotations with machine annotations and synthetic text images can yield accurate OCR for Sámi languages, even with a moderate amount of manually annotated data.

OCR萨米语文档数字化多语言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。