arXiv:2509.17768cs.CLcs.AI2025-09被引 3

评测主流语言识别模型在真实多语环境下的表现,发现其在混杂语言和非正式文本上严重失效。

DIVERS-Bench: Evaluating Language Identification Across Domain Shifts and Code-Switching

  • 构建跨领域、多语言混杂的综合评测基准DIVERS-BENCH
  • 模型在噪声和非正式文本上准确率显著下降,最高降幅达30%
  • 提出10种语言对的混杂文本数据集DIVERS-CS,适合评估真实场景中的语言识别能力

语言识别(LID)是多语言自然语言处理的核心任务,但现有系统常在干净、单一语言的数据上过拟合。本文提出DIVERS-BENCH,全面评估先进LID模型在多种领域的表现,包括语音转录、网页文本、社交媒体内容、儿童故事及语言混杂文本。结果显示,尽管模型在整理好的数据集上表现优异,但在嘈杂和非正式输入中性能急剧下降。我们还引入DIVERS-CS——一个覆盖10种语言对的多样化语言混杂基准数据集,表明现有模型难以识别同一句中的多种语言。这些发现凸显了在真实场景中构建更鲁棒、更具包容性的语言识别系统的重要性。

原文摘要 · Abstract (English)

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse domains, including speech transcripts, web text, social media texts, children's stories, and code-switched text. Our findings reveal that while models achieve high accuracy on curated datasets, performance degrades sharply on noisy and informal inputs. We also introduce DIVERS-CS, a diverse code-switching benchmark dataset spanning 10 language pairs, and show that existing models struggle to detect multiple languages within the same sentence. These results highlight the need for more robust and inclusive LID systems in real-world settings.

语言识别多语言评估基准代码混杂

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。