首个土耳其语语言最小对基准,评估大模型的语法理解能力
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
- 构建16类语法现象的土耳其语最小对数据集,每类1000对
- 顶尖大模型在人类易懂的语法任务上仍表现不佳
- 揭示模型对词序与形态复杂度的敏感性不同于人类
我们提出了TurBLiMP,首个针对土耳其语的语言最小对基准,用于评估单语和多语语言模型的语法能力。该基准涵盖16种语言现象,每类包含1000个最小对,填补了土耳其语语言评估资源的重要空白。设计时特别关注土耳其语中长期被忽视的两个特性:词序灵活性和通过形态过程实现的从属关系。我们在多种语言模型及新收集的人类可接受性判断数据上进行实验,发现即使最先进的大模型在人类轻松应对的语法任务上依然表现不足,且对词序和形态复杂度的敏感性与人类存在差异。
原文摘要 · Abstract (English)
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。