arXiv:2501.08322cs.CL2025-01被引 9

测试多语言大模型在真实拼写错误下的鲁棒性,发现性能下降2.3~4.3个百分点。

Exploring Robustness of Multilingual LLMs on Real-World Noisy Data

  • 基于维基编辑历史构建六种语言的真实拼写噪声词典。
  • 9个模型在含噪数据上平均性能下降2.3至4.3个百分点。
  • mT5系列模型表现最优,13B版本在多数语言中最具鲁棒性。

大型语言模型(LLMs)在包含人类拼写错误的网络数据上训练,但它们对真实世界中的类似噪声是否具备鲁棒性?本文研究了9个参数量从0.2B到13B不等的语言模型,在3项NLP任务(自然语言推理、命名实体识别、意图分类)中面对真实拼写错误的表现。实验覆盖6种语言,利用维基百科编辑历史构建真实噪声词典。结果显示,各模型在干净与含噪测试集间的性能差距,跨数据集和语言平均为2.3至4.3绝对百分点。总体而言,mT5模型比BLOOM、Falcon及BERT类模型更具鲁棒性;其中,mT5 (13B) 在所有任务和6种语言中的4种里表现最稳健。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are trained on Web data that might contain spelling errors made by humans. But do they become robust to similar real-world noise? In this paper, we investigate the effect of real-world spelling mistakes on the performance of 9 language models, with parameters ranging from 0.2B to 13B, in 3 different NLP tasks, namely Natural Language Inference (NLI), Name Entity Recognition (NER), and Intent Classification (IC). We perform our experiments on 6 different languages and build a dictionary of real-world noise for them using the Wikipedia edit history. We show that the performance gap of the studied models on the clean and noisy test data averaged across all the datasets and languages ranges from 2.3 to 4.3 absolute percentage points. In addition, mT5 models, in general, show more robustness compared to BLOOM, Falcon, and BERT-like models. In particular, mT5 (13B), was the most robust on average overall, across the 3 tasks, and in 4 of the 6 languages.

多语言鲁棒性拼写错误LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。