arXiv:2510.20957cs.CL2025-10被引 2

首个爱尔兰语语法能力评估基准,揭示大模型与人类差异。

Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting

  • 构建1020组最小对立对,覆盖11类语法特征,由母语者手工标注
  • 人类平均准确率90.1%,最强模型仅73.5%,差距达18.1%
  • 模型在不同语法点上表现不佳,适合低资源语言研究者参考

我们提出爱尔兰语最小对立对基准(Irish-BLiMP),首个面向濒危语言爱尔兰语的细粒度语言能力评估数据集与框架。基于多部语言学文献与语法参考书,由流利母语者团队手工构建并审核了1020组最小对立对,覆盖11类语法特征。我们评估了现有大型语言模型(LLMs)与母语者在爱尔兰语句法知识上的表现。结果显示,人类在所有语法特征上均优于所有模型,平均高出16.6%。开源自述模型与闭源模型间存在18.1%的显著差距,即使最强模型(gpt-5)准确率也仅达73.5%,远低于人类的90.1%。有趣的是,人类与模型在不同语法方面表现各异,表明模型学习到的表征机制存在偏差。Irish-BLiMP为评估爱尔兰语大模型语法能力提供了首个系统性框架,为低资源语言理解研究提供重要基准。

原文摘要 · Abstract (English)

We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing on a variety of linguistic literature and grammar reference works, we manually constructed and reviewed 1020 minimal pairs across a taxonomy of 11 linguistic features, through a team of fluent Irish speakers. We evaluate both existing Large Language Models (LLMs) and fluent human participants on their syntactic knowledge of Irish. Our findings show that humans outperform all models across all linguistic features, achieving 16.6% higher accuracy on average. Moreover, a substantial performance gap of 18.1% persists between open- and closed-source LLMs, with even the strongest model (gpt-5) reaching only 73.5% accuracy compared to 90.1% by human. Interestingly, human participants and models struggle on different aspects of Irish grammar, thus highlighting a difference in representation learned by the models. Overall, Irish-BLiMP provides the first systematic framework for evaluating the grammatical competence of LLMs in Irish and offers a valuable benchmark for advancing research on linguistic understanding in low-resource languages.

低资源语言语法评估大模型测评爱尔兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。