arXiv:2603.10861cs.CL2026-03中稿 · the 15th Language …被引 1

构建了斯里兰卡僧伽罗语历时语料库,覆盖1800至1955年共24万词

SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0

  • 基于国家图书馆文献,用OCR与人工校对构建语料
  • 含185部作品24.4万词,59篇标注了写作年代的7万词子集
  • 分虚构/非虚构与宗教、历史等细粒度类别,助力低资源语言研究

SiDiaC-v.2.0是迄今最大的僧伽罗语历时语料库,涵盖公元1800年至1955年的出版物,时间跨度从公元5世纪至20世纪。语料包含185部文学作品,总计24.4万词,经过严格筛选、预处理和版权合规检查,并进行大量后处理。其中59份文档(共7万词)按写作年代进行了标注。文本源自斯里兰卡国家图书馆,来自SiDiaC-v.1.0未过滤列表,使用Google Document AI OCR技术数字化。后续通过后处理修正格式、解决语码转换、添加特殊标记并修复异常标记。语料构建参考了FarPaHC、SiDiaC-v.1.0和CCOHA等语料库的经验,尤其在句法标注与文本规范化方面。语料分为两层:一级分类为虚构与非虚构;二级分类包括宗教、历史、诗歌、语言学与医学等具体体裁。尽管面临资源有限挑战,该语料库为僧伽罗语自然语言处理提供了全面支持,延续了SiDiaC-v.1.0的工作。

原文摘要 · Abstract (English)

SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the 5th to the 20th century CE in terms of written dates. The corpus consists of 244k words across 185 literary works that underwent thorough filtering, preprocessing, and copyright compliance checks, followed by extensive post-processing. Additionally, a subset of 59 documents totalling 70k words was annotated based on their written dates. Texts from the National Library of Sri Lanka were selected from the SiDiaC-v.1.0 non-filtered list, which was digitised using Google Document AI OCR. This was followed by post-processing to correct formatting issues, address code-mixing, include special tokens, and fix malformed tokens. The construction of SiDiaC-v.2.0 was informed by practices from other corpora, such as FarPaHC, SiDiaC-v.1.0, and CCOHA. This was particularly relevant for syntactic annotation and text normalisation strategies, given the shared characteristics of low-resource language status between Faroese and the similar cleaning strategies utilised in CCOHA. This corpus is categorised into two layers based on genres: primary and secondary. The primary categorisation is binary, assigning each book to either Non-Fiction or Fiction. The secondary categorisation is more detailed, grouping texts under specific genres such as Religious, History, Poetry, Language, and Medical. Despite facing challenges due to limited resources, SiDiaC-v.2.0 serves as a comprehensive resource for Sinhala NLP, building upon the work previously done in SiDiaC-v.1.0.

语料库僧伽罗语历时分析低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。