arXiv:2604.11066cs.CL2026-04被引 1

构建了迄今最大的克什米尔语预训练数据集,支持语言模型研究。

ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset

  • 从古籍文献和网络文本中收集并清洗出509万词的克什米尔语数据
  • 数据集含2760万字符、29.5万独特词型,子词平均每词2.38个
  • 适合克什米尔语自然语言处理研究者与模型开发者使用

我们提出KS-PRET-5M,目前公开可用的最大克什米尔语预训练数据集,包含5,090,244(509万)词、27,692,959(2760万)字符,以及295,433(29.5万)个唯一词型。数据源自两类来源:通过Malik~ ocite{malik2024inpage}转换器从专有InPage格式恢复的数字化档案与文学资料(涵盖文学、新闻、传记、小说、诗歌、宗教学术与学术写作),以及从克什米尔语网络资源获取的原生Unicode文本。所有文本经十一阶段清洗流程处理,实现平均克什米尔书写比例0.9965,全数据集仅残留146字符的天城文污染。采用google/muril-base-cased进行经验性分词,得到每词平均2.383个子词,总子词数约1213万,显著高于基于非克什米尔波斯-阿拉伯文字类比的先前估计。该数据集以连续文本流形式发布,遵循CC BY 4.0许可,支持克什米尔语语言模型预训练、分词器训练及计算语言学研究。

原文摘要 · Abstract (English)

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the dataset from two source classes: digitized archival and literary material, encompassing literature, news, biographies, novels, poetry, religious scholarship, and academic writing, recovered from the proprietary InPage desktop-publishing format using the converter of Malik~\cite{malik2024inpage}, and Unicode-native text collected from Kashmiri-language web sources. All text was processed through an eleven-stage cleaning pipeline that achieves a mean Kashmiri script ratio of 0.9965, reducing Devanagari contamination to 146 characters across the full dataset. We tokenized the dataset empirically using google/muril-base-cased, yielding a subword ratio of 2.383 tokens per word and a total of approximately 12.13 million subword tokens, substantially higher than prior estimates derived from non-Kashmiri Perso-Arabic analogues. KS-PRET-5M is released as a single continuous text stream under CC~BY~4.0 to support language model pretraining, tokenizer training, and computational linguistic research for Kashmiri.

克什米尔语数据集预训练自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。