构建魁北克法语最小对测试集,评估大模型语法理解能力。
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- 人工修改政府官网句子生成1761组最小对,每组由12名母语者标注
- 大模型语法能力随规模提升,但深层语义理解仍普遍失败
- 发现模型在魁北克法语上表现显著下降,顶尖模型仍具跨方言鲁棒性
本文提出魁北克法语语言最小对基准(QFrBLiMP),用于评估大语言模型对魁北克法语典型语法现象的掌握程度。QFrBLiMP包含1,761组最小对,每组标注20个语言现象(LPs)。这些最小对由从魁北克政府机构官方在线资源中提取的句子经人工修改生成。每组由12名魁北克法语母语者判断哪一句更符合语法,以此作为人类基准。我们通过比较不同模型在每组中对正确句赋予更高概率的比例,评估其表现。结果表明,尽管模型规模越大语法能力越强,但仍存在明显难度层级;所有模型在需深度语义理解的现象上持续失败,暴露出关键局限。统计分析显示,多数模型在魁北克法语上的性能显著低于通用法语(MultiBLiMP),但最先进模型仍在统计显著区间内,表明其具备一定的跨方言鲁棒性。
原文摘要 · Abstract (English)
In this paper, we introduce the Quebec-French Benchmark of Linguistic Minimal Pairs (QFrBLiMP), a corpus designed to evaluate LLMs' linguistic knowledge of prominent grammatical phenomena in Quebec-French. QFrBLiMP comprises 1,761 minimal pairs annotated with 20 LPs. Specifically, these minimal pairs have been created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution. Each pair is annotated by 12 Quebec-French native speakers, who select the sentence they consider grammatical from the two. These annotations are used to compare the competency of LLMs with that of humans. We evaluate different LLMs on QFrBLiMP and MultiBLiMP-Fr by observing the rate of higher probabilities assigned to the sentences of each minimal pair for each category. We find that while grammatical competence scales with model size, a clear hierarchy of difficulty emerges. All benchmarked models consistently fail on phenomena requiring deep semantic understanding, revealing a critical limitation. Finally, our statistical analysis comparing QFrBLiMP and MultiBLiMP reveals a significant performance degradation for most models on Quebec-French; however, the most capable models remain within the statistical significance interval, demonstrating cross-dialectal robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。