开源斯里兰卡多语种法律新闻政策文档集,支持多领域研究。
Sri Lanka Document Datasets: A Large-Scale, Multilingual Resource for Law, News, and Policy
- 构建26个数据集,涵盖议会、判决、政府文件等多类型文本。
- 含27.8万份文档,总大小80.7GB,覆盖僧伽罗语、泰米尔语和英语。
- 每日更新,开源可下载,适合语言处理与社会政治研究者使用。
我们发布一组开放、机器可读的斯里兰卡文档数据集,涵盖议会记录、法律判决、政府出版物、新闻及旅游统计数据。当前共包含26个数据集,总计278,621份文档(80.7 GB),语种包括僧伽罗语、泰米尔语和英语。数据每日更新,并同步至GitHub与Hugging Face。这些资源旨在支持计算语言学、法律分析、社会政治研究及多语言自然语言处理等领域研究。本文介绍数据来源、采集流程、格式规范及潜在应用场景,同时讨论版权与伦理问题。本稿件版本为v2026-07-02-0940。
原文摘要 · Abstract (English)
We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 278,621 documents (80.7 GB) across 26 datasets in Sinhala, Tamil, and English. The datasets are updated daily and mirrored on GitHub and Hugging Face. These resources aim to support research in computational linguistics, legal analytics, socio-political studies, and multilingual natural language processing. We describe the data sources, collection pipeline, formats, and potential use cases, while discussing licensing and ethical considerations. This manuscript is at version v2026-07-02-0940.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。