arXiv:2603.04854cs.CL2026-03中稿 · the 2nd workshop o…被引 1

首个斯里兰卡立法文本语料库,助力僧伽罗语法律NLP研究。

SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts

  • 构建200万词的僧伽罗语立法文本数据集,含1065部法案与141份议案。
  • 通过OCR与人工校对确保高质量,支持法律信息抽取与分析任务。
  • 适用于法律NLP、语言模型评估及僧伽罗语资源建设的研究者。

SinhaLegal引入了一个包含约200万词的僧伽罗语立法文本语料库,涵盖1,206份法律文件,包括1,065部1981至2014年间的法案和141份2010至2014年的议案,均来自官方渠道系统收集。文本通过Google Document AI进行光学字符识别(OCR),并经过大量后处理与人工清理,确保内容可机器读取,每份文档附有专用元数据文件。全面评估包括语料库统计、词汇多样性、词频分析、命名实体识别及主题建模,证实该语料库具有结构化与领域特异性。此外,使用大模型与小模型进行了困惑度分析,评估语言模型在领域文本上的表现。SinhaLegal是支持摘要生成、信息抽取与分析等自然语言处理任务的重要资源,填补了僧伽罗语法律研究中的关键空白。

原文摘要 · Abstract (English)

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to 2014, which were systematically collected from official sources. The texts were extracted using OCR with Google Document AI, followed by extensive post-processing and manual cleaning to ensure high-quality, machine-readable content, along with dedicated metadata files for each document. A comprehensive evaluation was conducted, including corpus statistics, lexical diversity, word frequency analysis, named entity recognition, and topic modelling, demonstrating the structured and domain-specific nature of the corpus. Additionally, perplexity analysis using both large and small language models was performed to assess how effectively language models respond to domain-specific texts. The SinhaLegal corpus represents a vital resource designed to support NLP tasks such as summarisation, information extraction, and analysis, thereby bridging a critical gap in Sinhala legal research.

法律NLP语料库僧伽罗语信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。