arXiv:2508.01006cs.CL2025-08ACL被引 2

构建首个乌尔都语语法能力评测基准,检验大模型对细微语法差异的识别能力。

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

  • 设计5696组语义极近但语法正误不同的句子对,覆盖10类核心句法现象。
  • 顶尖模型如LLaMA-3-70B平均准确率达94.73%,但各模型表现差异显著。
  • 适合评估低资源语言模型的语法理解能力,尤其关注乌尔都语等小语种。

多语言大语言模型在多种语言上表现出色,但在低资源语言如乌尔都语上的训练数据远少于英语等高资源语言。为评估大模型在乌尔都语中的语言知识,本文提出乌尔都语语言学最小对立对基准(UrBLiMP),包含5,696组仅在语法可接受性上存在细微差异的句子对,覆盖十类核心句法现象,基于乌尔都语树库和多样文本语料精心构建。人工标注的跨标注者一致性达96.10%,验证了数据集可靠性。我们对20个主流多语言大模型进行了评估,结果显示不同模型在各类句法现象上的表现差异显著。尽管LLaMA-3-70B取得最高平均准确率(94.73%),其性能与Gemma-3-27B-PT等其他顶级模型无统计显著差异。这些发现揭示了当前多语言大模型在捕捉低资源语言精细句法知识方面的潜力与局限。

原文摘要 · Abstract (English)

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to high-resource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages.

语言模型乌尔都语评测基准句法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。