arXiv:2601.11170cs.CL2026-01中稿 · the LREC 2026 conf…

构建南斯拉夫语族网页语料库,两年迭代后内容新旧重叠仅20%。

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

  • 通过持续抓取国家域名,实现多语言语料自动采集。
  • 产出7种语言共170亿词、3810万文本的大型语料库。
  • 发现机器生成内容增多导致网页质量下降,需警惕数据污染。

抓取国家顶级域名已被证明是获取低资源语言文本的有效方法。该方法近期被应用于南斯拉夫语言,产生了该语言群体最大的通用语料库——CLASSLA-web 1.0。在此基础上,我们建立了针对南斯拉夫及关联网络的连续性国家顶级域名爬取基础设施。本文呈现该系统的首个成果:CLASSLA-web 2.0语料库,包含170亿词、3810万文本,覆盖波斯尼亚语、保加利亚语、克罗地亚语、马其顿语、黑山语、塞尔维亚语和斯洛文尼亚语七种语言。除文体分类外,新版本还自动标注主题标签。对比显示,与前代语料仅20%文本重叠,表明两年后重新爬取已获得大量新内容。然而,尽管收益增加,也出现新问题:对主要域名的人工检查发现网页内容明显退化,机器生成网站已占相当比例。

原文摘要 · Abstract (English)

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general corpora for this language group: the CLASSLA-web 1.0 corpora. Building on this success, we established a continuous crawling infrastructure for iterative national top-level domain crawling across South Slavic and related webs. We present the first outcome of this crawling infrastructure - the CLASSLA-web 2.0 corpus collection, with substantially larger web corpora containing 17.0 billion words in 38.1 million texts in seven languages: Bosnian, Bulgarian, Croatian, Macedonian, Montenegrin, Serbian, and Slovenian. In addition to genre categories, the new version is also automatically annotated with topic labels. Comparing CLASSLA-web 2.0 with its predecessor reveals that only one-fifth of the texts overlap, showing that re-crawling after just two years yields largely new content. However, while the new web crawls bring growing gains, we also notice growing pains - a manual inspection of top domains reveals a visible degradation of web content, as machine-generated sites now contribute a significant portion of texts.

语料库自然语言处理网页爬取南斯拉夫语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。