arXiv:2601.01091cs.CLcs.AI2026-01被引 1

构建310万词克什米尔语数据集,助力大模型理解该语言。

ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining

  • 开发专用转换工具,将老旧排版格式文献转为可训练文本。
  • 数据集含310万词、1640万字符,覆盖文学、新闻等多元文体。
  • 适合研究克什米尔语NLP或低资源语言模型的学者使用。

大型语言模型在高资源语言上表现优异,但在克什米尔语(约七百万使用者)中却难以生成连贯文本。这一差距并非模型本身缺陷,而是高质量训练数据极度匮乏所致。由于数十年来的克什米尔语文献多以专有InPage格式编码,无法被现代自然语言处理系统访问。本文提出KS-LIT-3M,一个包含310万词(1640万字符)的克什米尔语语料库,专为语言模型预训练设计。数据以连续线性文本流形式组织,适用于因果语言模型训练。通过自主研发的InPage-to-Unicode转换器,并经英文污染清除、字符归一化与质量验证等严格预处理流程构建。语料涵盖131,607个独特词汇,内容来自文学作品、新闻报道、学术论文及宗教文献等多种体裁,填补了克什米尔语技术资源空白。数据集以CC-BY-4.0许可开放,推动克什米尔语自然语言处理研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approximately seven million people. This performance disparity stems not from inherent model limitations but from a critical scarcity of high-quality training data. Decades of Kashmiri literature remain inaccessible to modern NLP pipelines due to their encoding in the proprietary InPage desktop publishing format. This paper introduces KS-LIT-3M, a curated corpus of 3.1 million words (16.4 million characters) specifically designed for pretraining language models on Kashmiri. The dataset is structured as a single continuous linear text stream, optimized for causal language model training where models learn to predict subsequent tokens from preceding context. The corpus was constructed through the development of a specialized InPage-to-Unicode converter, followed by rigorous preprocessing including English contamination removal, character normalization, and quality validation. Encompassing 131,607 unique words drawn from diverse genres including literary works, journalistic writing, academic texts, and religious scholarship, KS-LIT-3M addresses a fundamental resource gap for Kashmiri language technology. The dataset is released under the CC-BY-4.0 license to facilitate research in Kashmiri natural language processing.

克什米尔语数据集语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。