arXiv:2502.06692cs.CLcs.AI2025-02被引 3

构建多标签北欧语言识别数据集,提升句子级多语言识别准确率

Multi-label Scandinavian Language Identification (SLIDE)

  • 提出多标签训练方法,同时识别丹麦语、挪威语等四种北欧语言
  • 在新构建的SLIDE数据集上,多标签模型准确率显著优于单标签模型
  • 适合需要处理混合语言文本的自然语言处理研究者使用

句子级语言识别在相近语言间尤为困难,常需多语言标注。本文聚焦丹麦语、挪威语博克马尔体、挪威语尼诺斯克体和瑞典语的多标签句子级语言识别(LID)。我们提出了手动标注的多标签评估数据集SLIDE及一系列具有不同速度-精度权衡的LID模型。实验表明,同时识别多种语言对实现高精度至关重要,并提出一种新型多标签训练方法。

原文摘要 · Abstract (English)

Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focus on multi-label sentence-level Scandinavian language identification (LID) for Danish, Norwegian Bokmål, Norwegian Nynorsk, and Swedish. We present the Scandinavian Language Identification and Evaluation, SLIDE, a manually curated multi-label evaluation dataset and a suite of LID models with varying speed-accuracy tradeoffs. We demonstrate that the ability to identify multiple languages simultaneously is necessary for any accurate LID method, and present a novel approach to training such multi-label LID models.

语言识别多标签北欧语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。