评测大模型对巴斯克语和西班牙语方言的理解能力,发现方言差异导致性能下降。
Lost in Variation? Evaluating NLI Performance in Basque and Spanish Geographical Variants
- 构建巴斯克语与西班牙语及其方言的平行数据集用于自然语言推理任务
- 编码器模型在处理西部巴斯克语时表现显著下降,尤其在跨方言场景下
- 研究揭示方言差异本身是性能下降主因,而非词汇重叠问题
本文评估当前语言技术对巴斯克语和西班牙语语言变体的理解能力。以自然语言推理(NLI)为基准任务,构建了一个新的人工标注的巴斯克语与西班牙语平行数据集及其各自方言版本。基于编码器型与解码器型大语言模型的跨语言及上下文学习实验表明,面对语言变异时性能明显下降,尤以巴斯克语为甚。错误分析显示,该下降并非由词汇重叠引起,而是源于语言变异本身。进一步消融实验表明,编码器模型在处理西部巴斯克语时尤为困难,这与语言学理论一致——边缘方言(如西部)相较于标准语距离更远。所有数据与代码均已公开。
原文摘要 · Abstract (English)
In this paper, we evaluate the capacity of current language technologies to understand Basque and Spanish language varieties. We use Natural Language Inference (NLI) as a pivot task and introduce a novel, manually-curated parallel dataset in Basque and Spanish, along with their respective variants. Our empirical analysis of crosslingual and in-context learning experiments using encoder-only and decoder-based Large Language Models (LLMs) shows a performance drop when handling linguistic variation, especially in Basque. Error analysis suggests that this decline is not due to lexical overlap, but rather to the linguistic variation itself. Further ablation experiments indicate that encoder-only models particularly struggle with Western Basque, which aligns with linguistic theory that identifies peripheral dialects (e.g., Western) as more distant from the standard. All data and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。