对比六种OCR引擎在僧伽罗语和泰米尔语上的零样本识别效果。
Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil
- 对比商用与开源OCR在两种低资源语言上的表现。
- Surya对僧伽罗语识别准确率最高(词错误率2.61%)。
- 首次构建合成泰米尔语OCR测试数据集,适合多语言研究者。
针对使用独特文字系统的低资源语言(LRL),光学字符识别(OCR)仍是未解难题。本文对比分析了六种不同OCR引擎在僧伽罗语和泰米尔语上的零样本性能。评估系统包括Cloud Vision API、Surya、Document AI、Tesseract(均支持双语言)、Subasa OCR与EasyOCR(仅支持单语言)。通过五种测量方法,在字符与词级层面评估准确率。结果显示,Surya在僧伽罗语上表现最佳,词错误率(WER)为2.61%;Document AI在泰米尔语上最优,字符错误率(CER)低至0.78%。此外,本文还构建了一个新型合成泰米尔语OCR基准数据集。
原文摘要 · Abstract (English)
Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Languages (HRL). However, for Low-Resourced Languages (LRL) that use unique scripts, it remains an open problem. This study presents a comparative analysis of the zero-shot performance of six distinct OCR engines on two LRLs: Sinhala and Tamil. The selected engines include both commercial and open-source systems, aiming to evaluate the strengths of each category. The Cloud Vision API, Surya, Document AI, and Tesseract were evaluated for both Sinhala and Tamil, while Subasa OCR and EasyOCR were examined for only one language due to their limitations. The performance of these systems was rigorously analysed using five measurement techniques to assess accuracy at both the character and word levels. According to the findings, Surya delivered the best performance for Sinhala across all metrics, with a WER of 2.61%. Conversely, Document AI excelled across all metrics for Tamil, highlighted by a very low CER of 0.78%. In addition to the above analysis, we also introduce a novel synthetic Tamil OCR benchmarking dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。