arXiv:2601.01088cs.CVcs.CL2026-01

构建60万张克什米尔语文字图像数据集,助力濒危语言识别

600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script

  • 用合成方法生成60万张克什米尔语文字图像,适配多种识别模型
  • 图像分辨率256x64,含真实文档退化模拟与背景纹理增强鲁棒性
  • 专为低资源语言OCR设计,适合研究濒危语言数字化的学者

本文介绍600K-KS-OCR数据集,一个大规模合成语料库,包含约60.2万张单词级分割图像,用于训练和评估针对克什米尔语的光学字符识别系统。该语言属于使用改良波斯-阿拉伯字母的达迪克语系,约有七百万使用者,现处于濒危状态。每张图像尺寸为256x64像素,附带多格式真实转录文本,兼容CRNN、TrOCR及通用机器学习流程。生成方法融合三种传统克什米尔字体,通过数据增强模拟真实文档退化,并引入多样背景纹理以提升模型鲁棒性。数据集分为十个分区,总大小约10.6 GB,采用CC-BY-4.0许可发布,旨在推动低资源语言光学字符识别研究。

原文摘要 · Abstract (English)

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising approximately 602,000 word-level segmented images designed for training and evaluating optical character recognition systems targeting Kashmiri script. The dataset addresses a critical resource gap for Kashmiri, an endangered Dardic language utilizing a modified Perso-Arabic writing system spoken by approximately seven million people. Each image is rendered at 256x64 pixels with corresponding ground-truth transcriptions provided in multiple formats compatible with CRNN, TrOCR, and generalpurpose machine learning pipelines. The generation methodology incorporates three traditional Kashmiri typefaces, comprehensive data augmentation simulating real-world document degradation, and diverse background textures to enhance model robustness. The dataset is distributed across ten partitioned archives totaling approximately 10.6 GB and is released under the CC-BY-4.0 license to facilitate research in low-resource language optical character recognition.

OCR克什米尔语合成数据低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。