arXiv:2602.13139cs.CL2026-02被引 4

提升近义语言识别精度,解决低资源语言噪声问题

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report

  • 扩充训练数据并合并相似语言簇,引入噪声标签
  • 在三组近义语言上精度显著提升,尤其对低资源语言
  • 适合构建高质量多语言语料库的研究者使用

语言识别(LID)是利用网络数据构建高质量多语言数据集的关键步骤。现有工具(如OpenLID或GlotLID)难以区分密切相关的语言,也难分辨真实语言与噪声,导致语言子集污染,尤其影响低资源语言。本文通过扩充训练数据、合并问题语言变体簇、引入噪声标记标签,升级OpenLID为OpenLID-v3。在多个基准测试中对比GlotLID,重点评估波斯尼亚、克罗地亚、塞尔维亚语;意大利北部与法国南部罗曼语方言;以及北欧语言三组近义语言。针对现有数据集不足的情况,贡献了新的评估数据集。实验发现集成方法虽提升精度,但大幅降低低资源语言覆盖率。OpenLID-v3已发布于https://huggingface.co/HPLT/OpenLID-v3。

原文摘要 · Abstract (English)

Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid natural language from noise, which contaminates language-specific subsets, especially for low-resource languages. In this work we extend the OpenLID classifier by adding more training data, merging problematic language variant clusters, and introducing a special label for marking noise. We call this extended system OpenLID-v3 and evaluate it against GlotLID on multiple benchmarks. During development, we focus on three groups of closely related languages (Bosnian, Croatian, and Serbian; Romance varieties of Northern Italy and Southern France; and Scandinavian languages) and contribute new evaluation datasets where existing ones are inadequate. We find that ensemble approaches improve precision but also substantially reduce coverage for low-resource languages. OpenLID-v3 is available on https://huggingface.co/HPLT/OpenLID-v3.

语言识别低资源语言噪声过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。