arXiv:2605.18232cs.CLcs.AI2026-05被引 1

首个高质量索马里语语料库,含专用分词器与语言识别评测基准。

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark

论文配图:SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
图 1 · 摘自论文原文
  • 从三大源数据构建6阶段可复现语料过滤流程
  • 产出81万文档约3.03亿词元的高质量语料
  • 提供首个公开的索马里语语言识别对比评测

索马里语是非洲之角的库希特语系,拥有约2500万使用者,但此前尚未有公开发布的专用预训练语料库、配套分词器及语言识别评测基准。现有索马里语文本要么包含在多语种数据集(如HPLT v2、CC100、MADLAD-400、OSCAR、mC4)中,要么以小型未标注的索马里语数据集形式出现在Hugging Face上。本文提出SomaliWeb v1,基于HPLT v2、CC100和索马里维基百科三个上游来源,通过六阶段可复现的清洗流程,构建了一个包含819,322篇文档(约3.03亿词元)的质量筛选语料库。我们公开发布:(i) 语料库,(ii) 匹配的BPE-16K分词器,以及 (iii) 首个公开的三款主流语言识别器并行对比评测。测量显示,现有数据存在明显质量问题:HPLT v2的“清理后”索马里语版本仍保留17.3%的字节完全重复项,56.1%的文档存在可修复的摩吉巴克(mojibake),10.7%的字节唯一文档在Jaccard tau=0.80条件下为近似重复。我们的BPE-16K分词器在FLORES-200索马里语开发测试集上的词元数比GPT-4的cl100k_base少40.2%,作为分词器层面的量化指标;下游语言模型困惑度比较将在后续版本中发布。

原文摘要 · Abstract (English)

Somali is a Cushitic language of the Horn of Africa with ~25 million speakers, yet no documented dedicated Somali pretraining corpus with a companion tokenizer and language-identification benchmark has been publicly released. Existing Somali text appears either inside multilingual distributions (HPLT v2, CC100, MADLAD-400, OSCAR, mC4) or in small, undocumented Somali-only uploads on Hugging Face. We introduce SomaliWeb v1, a quality-filtered Somali corpus of 819,322 documents (~303M tokens) built from three upstream sources (HPLT v2, CC100, Somali Wikipedia) through a six-stage reproducible pipeline. We release (i) the corpus, (ii) a matched BPE-16K tokenizer, and (iii) the first public side-by-side Somali benchmark of three production language identifiers. Our measurements reveal concrete quality defects in existing distributions: HPLT v2's "cleaned" Somali release retains 17.3% byte-exact duplicates, 56.1% of its documents contain fixable mojibake, and 10.7% of its byte-unique documents are near-duplicates at Jaccard tau=0.80. Our BPE-16K tokenizer emits 40.2% fewer tokens than GPT-4's cl100k_base on FLORES-200 Somali devtest as a tokenizer-level measurement; downstream language-model perplexity comparisons are deferred to a follow-up release.

语料库索马里语分词器语言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。