构建首个覆盖12项任务的尼泊尔语NLU基准,推动低资源语言研究
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
- 新增12个尼泊尔语任务,涵盖分类、相似性、自然语言推理等
- 多语言模型在多数任务上优于单语模型,凸显跨语言优势
- 为尼泊尔语等低资源语言提供可复现的评估标准
尼泊尔语具有独特的语言特征,如复杂的德纳加里文字、形态学和多种方言,给自然语言理解(NLU)任务带来独特挑战。尽管已有尼泊尔语言理解评估(Nep-gLUE)基准,但其仅涵盖四项任务,难以全面评估自然语言处理(NLP)模型性能。为此,我们引入十二个新数据集,构建新的尼泊尔语语言理解评估(NLUE)基准,用于评估模型在多样化自然语言理解任务上的表现。新增任务包括单句分类、相似性与改写任务、自然语言推断(NLI)以及通用掩码评估任务(GMET)。通过大量实验发现,现有顶级模型在这些新任务上表现不佳。同时,最佳多语言模型在多数任务上优于最佳单语模型,表明需要更鲁棒的解决方案来应对尼泊尔语复杂性。该扩展基准为评估、比较和推进模型提供了新标准,显著促进低资源语言的NLP研究。
原文摘要 · Abstract (English)
The Nepali language has distinct linguistic features, especially its complex script (Devanagari script), morphology, and various dialects,which pose a unique challenge for Natural Language Understanding (NLU) tasks. While the Nepali Language Understanding Evaluation (Nep-gLUE) benchmark provides a foundation for evaluating models, it remains limited in scope, covering four tasks. This restricts their utility for comprehensive assessments of Natural Language Processing (NLP) models. To address this limitation, we introduce twelve new datasets, creating a new benchmark, the Nepali /Language Understanding Evaluation (NLUE) benchmark for evaluating the performance of models across a diverse set of Natural Language Understanding (NLU) tasks. The added tasks include Single-Sentence Classification, Similarity and Paraphrase Tasks, Natural Language Inference (NLI), and General Masked Evaluation Task (GMET). Through extensive experiments, we demonstrate that existing top models struggle with the added complexity of these tasks. We also find that the best multilingual model outperforms the best monolingual models across most tasks, highlighting the need for more robust solutions tailored to the Nepali language. This expanded benchmark sets a new standard for evaluating, comparing, and advancing models, contributing significantly to the broader goal of advancing NLP research for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。