arXiv:2604.02135cs.CL2026-04被引 1

首个苏格兰盖尔语多维度评测基准,揭示大模型在语法上已超人类水平

GaelEval: Benchmarking LLM Performance for Scottish Gaelic

  • 构建三类任务:语法选择题、文化翻译题和知识问答题
  • 谷歌Gemini 3 Pro在语法任务中达83.3%准确率,超过人类78.1%
  • 专有模型普遍优于开源模型,用盖尔语提示提升2.4%表现

多语言大模型常在未官方支持的语言中表现出隐性能力,但其在这些语言上的表现参差不齐且缺乏衡量。这一问题在形态句法复杂的少数族裔语言如苏格兰盖尔语中尤为突出,现有翻译基准无法捕捉其结构能力。本文提出GaelEval,首个针对盖尔语的多维度评测基准,包含:(i) 专家编写的形态句法多选题;(ii) 贴近文化的翻译任务;(iii) 大规模文化知识问答任务。对19个LLM进行评估,与30位流利母语者基准对比,发现Gemini 3 Pro Preview在语言任务中达到83.3%准确率,高于人类基准(78.1%)。专有模型整体优于开源模型,使用盖尔语提示带来+2.4%的稳定提升。文化任务中,领先模型准确率超90%,但多数系统在盖尔语提示下表现下降,且绝对分数因人工基准而被高估。总体表明,前沿模型在盖尔语多个语法维度已超越人类,验证了盖尔语提示的有效性,并揭示专有与开源模型间的持续性能差距。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) often exhibit emergent 'shadow' capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline ($n=30$), we find that Gemini 3 Pro Preview achieves $83.3\%$ accuracy on the linguistic task, surpassing the human baseline ($78.1\%$). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+$2.4\%$). On the cultural task, leading models exceed $90\%$ accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.

语言评测小语种大模型盖尔语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。