arXiv:2604.25384cs.CL2026-04

将维基原始数据转化为七种南斯拉夫语言的高质量语料库。

Wiki Dumps to Training Corpora: South Slavic Case

论文配图:Wiki Dumps to Training Corpora: South Slavic Case
图 1 · 摘自论文原文
  • 从维基百科等源提取文本并清洗,保留自然语言内容。
  • 用n-gram方法识别重复内容,移除低质条目以提升质量。
  • 成果可支持语言模型训练与跨语言比较研究。

本文提出一个流程,将维基媒体原始数据转换为七种南斯拉夫语言的高质量文本语料库。整个流程分为两个阶段:第一阶段从维基百科、维基文库、维基教科书、维基新闻和维基语录的原始转储中提取并清洗文本,需妥善处理原始维基标记,分离出文章内容及可用的自然语言文本;第二阶段针对由数据库或知识库生成的低质条目(具有重复模式、通用表述、无原创内容)进行过滤,采用基于n-gram的冗余检测策略,识别并移除高重复性文章。最终构建的语料库旨在提供语言学信息丰富的文本,适用于语言模型训练或南斯拉夫语言间的比较研究。尽管聚焦南斯拉夫语族,该方法具备高度语言无关性,可推广至其他语言。

原文摘要 · Abstract (English)

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw dumps of Wikipedia, Wikisource, Wikibooks, Wikinews, and Wikiquote. This step requires careful handling of raw wiki markup to isolate, first of all, textual articles, and then usable natural language text within them. The second phase addresses the challenge of questionable or low-quality articles, which are often generated from databases or structured knowledge bases. These articles are characterised by repetitive patterns, generic phrasing, and minimal to no original content. To mitigate their impact, a n-gram-based filtering strategy was employed to detect high levels of textual redundancy between articles and then remove such articles from the corpora entirely. The resulting datasets aim to provide linguistically rich texts suitable for training language models or conducting comparative research across South Slavic languages. By combining systematic extraction with quality control, this work contributes to the creation of reliable, high-information corpora that reflect the authentic cultural contexts of languages. While focused on the South Slavic case in the paper, the approach is mostly language-agnostic and can be generalised to other languages.

语料库构建南斯拉夫语言文本清洗n-gram过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。