arXiv:2606.23144cs.CVcs.CL2026-06

首个克什米尔语大规模合成OCR数据集,助力低资源语言文字数字化。

Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri

  • 用SynthOCR-Gen生成61万+图像-文本对,覆盖多种字体与文档粒度。
  • 包含25种增强策略,模拟真实文档退化,提升模型鲁棒性。
  • 为克什米尔语OCR训练提供低成本高扩展资源,适合语言技术研究者。

光学字符识别(OCR)在低资源语言中常受限于标注数据匮乏和书写系统特有的渲染复杂性。克什米尔语主要使用波斯阿拉伯文纳斯塔利克书写,因上下文字形变化、密集连写及拼写变体而更具挑战。我们提出Koshur Pixel,首个针对克什米尔语的大规模合成OCR数据集,基于KS-PRET-5M语料库,利用SynthOCR-Gen框架生成613,078张图像-文本配对。数据集涵盖多种字体和文本粒度,从单个词到整页文档,并集成超过25种增强策略,以模拟真实文档退化。该数据集为训练OCR系统、数字化克什米尔文本遗产以及推动严重低资源语言的技术发展提供了可扩展且成本可控的基础资源。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-specific rendering. Kashmiri, written primarily in the Perso-Arabic Nastaliq script, presents additional challenges due to contextual glyph shaping, dense ligatures, and orthographic variability. We introduce Koshur Pixel, the first large-scale synthetic OCR dataset for Kashmiri, comprising 613,078 image-text pairs generated from the KS-PRET-5M corpus using the SynthOCR-Gen framework. The dataset spans multiple fonts and textual granularities, ranging from individual words to full-page documents, and incorporates more than 25 augmentation strategies that emulate real-world document degradations. Koshur Pixel provides a scalable and cost-effective alternative to manual annotation, establishing a foundational resource for training OCR systems, digitizing Kashmiri textual heritage, and advancing language technologies for a severely under-resourced language.

OCR合成数据低资源语言克什米尔语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。