构建109种语言的网页文本语言识别基准,揭示现有模型高估性能问题。
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
- 基于社区标注构建网页领域多语言识别基准
- 109种语言测试显示多数模型在真实场景下准确率偏低
- 适合关注低资源语言与真实数据评估的研究者
语言识别(LID)是构建多语言语料库的基础步骤。然而,现有LID模型在噪声大、异构性强的网络数据上表现仍不佳,尤其对许多低资源语言。本文提出CommonLID,一个面向网络领域的、由人工标注的109种语言的公共基准。该数据集覆盖了大量此前被忽视的语言,是构建更具代表性的高质量文本语料的关键资源。我们结合五个常见评估集,测试了八种主流LID模型,并分析结果以定位当前技术状态。特别指出,现有评估在网页场景中普遍高估了多种语言的识别准确率。我们已将CommonLID及构建代码以开放许可发布。
原文摘要 · Abstract (English)
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID's value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。