构建101种语言的语法最小对数据集,测试大模型跨语言能力
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
- 用自动化流程基于通用依存和UniMorph构建
- 覆盖101语言,超12.8万组语法最小对
- 揭示大模型在低资源语言上的短板
我们提出MultiBLiMP 1.0,一个大规模多语言语法最小对基准,涵盖101种语言和两种主谓一致类型,包含超过12.8万个最小对。这些最小对通过完全自动化的流水线生成,利用了Universal Dependencies和UniMorph的大规模语言学资源。MultiBLiMP 1.0以前所未有的多语言规模评估大语言模型的能力,并凸显了当前最先进模型在建模低资源语言方面的不足。
原文摘要 · Abstract (English)
We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。