arXiv:2606.29378cs.CL2026-06中稿 · the 12th Moratuwa …

首个真实世界斯里兰卡僧伽罗文版面级OCR数据集,验证了先进模型在历史文档上的强鲁棒性。

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

  • 构建首个真实印刷文本的僧伽罗文版面级数据集,覆盖1981-2019年立法文件
  • 轻量级模型LightOnOCR-2-1B在测试集上达1.05%字符错误率,显著优于主流开源与商用模型
  • 首次实现跨时间跨度的真实文档识别,适合历史文献数字化与语言技术研究者

僧伽罗语是一种在斯里兰卡有约1600万人使用的形态丰富的音节文字,但至今缺乏公开的真实世界版面级僧伽罗文OCR数据集。此前评估均依赖人工合成数据。为填补空白,我们推出了sinhala-ocr-lk-acts-1010数据集,包含1010张来自1981–1989及2000–2019年间斯里兰卡立法法案的页面图像及其转录文本,划分为707个训练样本、101个验证样本和202个测试样本。基于深度学习视觉语言模型,我们对DeepSeek-OCR V1、DeepSeek-OCR V2和LightOnOCR-2-1B三种模型采用QLoRA进行微调,在消费级与云端GPU上完成8次实验。LightOnOCR-2-1B表现最佳,在所有测试样本上达到1.05%的字符错误率(CER),显著优于Surya-OCR(8.84%)和Tesseract v5(10.69%)等开源模型,以及Google Document AI(2.06%)等商业模型。结果表明,该模型在真实世界任务中表现优异,并在不同印刷时期间保持稳定性能,即使在严重退化的文档上亦具鲁棒性。

原文摘要 · Abstract (English)

Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data. To bridge the gap, we introduce sinhala-ocr-lk-acts-1010, an annotated dataset of 1,010 page-level images and their transcriptions collected from Sri Lankan Legislative Acts published between 1981-1989 and 2000-2019, split into 707 training examples, 101 validation examples, and 202 testing examples. Three models based on deep learning-based visual language processing, namely DeepSeek-OCR V1, DeepSeek-OCR V2, and LightOnOCR-2-1B, are fine-tuned using QLoRA in 8 experiments conducted on consumer and cloud GPUs. LightOnOCR-2-1B is the top performer, achieving a CER of 1.05% across all test examples, outperforming state-of-the-art open-source OCR models such as Surya-OCR (8.84%) and Tesseract v5 (10.69%), as well as commercially available OCR models such as Google Document AI (2.06%). Our results suggest that LightOnOCR-2-1B outperforms other baselines on real-world OCR tasks and maintains consistent performance across all print periods, even when documents are severely degraded.

OCR历史文献多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。