arXiv:2603.29221cs.CL2026-03中稿 · the 2nd Workshop o…被引 1

构建了涵盖78万句的斯里兰卡佛教文本语料库,助力文化遗产数字化。

SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali

  • 采用高精度OCR与网络爬取,整合历史文献与三藏经文
  • 包含925万词,16份受版权许可文献,支持语言模型训练
  • 适合佛教研究、语言学分析及文化数字化保护者使用

SiPaKosa 是一个涵盖僧伽罗语和巴利语教义文本的综合性语料库,包含约78.6万句子和925万词,整合了16份获版权许可的历史佛教文献以及完整的网络爬取三藏经文。语料库通过 Google Document AI 对历史手稿进行高质量光学字符识别,并结合系统性网络爬取规范典籍资源,再经严格质量控制与元数据标注。语料库按语言划分为僧伽罗语与混合僧伽罗-巴利语子语料库。我们使用十种预训练模型评估其表现,困惑度范围为1.09至189.67,结果显示专有模型性能优于开源模型3到6倍。该语料库可用于领域适配语言模型的预训练、历史语言分析及佛教学术信息检索系统开发,同时致力于保护僧伽罗文化传承。

原文摘要 · Abstract (English)

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-scraped Tripitaka canonical texts. The corpus was created through high-quality OCR using Google Document AI on historical manuscripts, combined with systematic web scraping of canonical repositories, followed by rigorous quality control and metadata annotation. The corpus is organised into language-specific subcorpora: Sinhala and Mixed Sinhala-Pali. We evaluate the performance of language models using ten pretrained models, with perplexity scores ranging from 1.09 to 189.67 on our corpus. This analysis shows that proprietary models significantly outperform open-source alternatives by factors of three to six times. This corpus supports the pretraining of domain-adapted language models, facilitates historical language analysis, and aids in the development of information retrieval systems for Buddhist scholarship while preserving Sinhala cultural heritage.

语料库佛教文献文化遗产自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。