构建首个沙特方言多维度评测集,检验大模型语言文化理解能力
From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher
- 设计涵盖6类任务的1.8万道题评测集,覆盖沙特主要方言
- 多模型测试发现文化推理与语境理解存在显著性能短板
- 为提升真实场景阿拉伯语应用效果提供评估与训练方向
随着大语言模型在阿拉伯语自然语言处理中的日益重要,评估其对区域方言与文化细节的理解能力至关重要,尤其是在语言多样性突出的沙特阿拉伯。本文提出Absher,一个专门用于评估大模型在主要沙特方言中表现的综合性基准。该基准包含超过18,000道多项选择题,涵盖意义理解、真假判断、填空、语境使用、文化解读和地点识别六类任务,题目源自沙特各地搜集的方言词汇、短语与谚语。我们评估了多种前沿大模型,包括多语言及专精阿拉伯语的模型,并深入分析其能力与局限。结果揭示,在需要文化推断或语境理解的任务中存在明显性能差距。研究强调亟需开展方言感知训练与文化对齐的评估方法,以提升大模型在真实阿拉伯语应用中的表现。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly central to Arabic NLP applications, evaluating their understanding of regional dialects and cultural nuances is essential, particularly in linguistically diverse settings like Saudi Arabia. This paper introduces Absher, a comprehensive benchmark specifically designed to assess LLMs performance across major Saudi dialects. \texttt{Absher} comprises over 18,000 multiple-choice questions spanning six distinct categories: Meaning, True/False, Fill-in-the-Blank, Contextual Usage, Cultural Interpretation, and Location Recognition. These questions are derived from a curated dataset of dialectal words, phrases, and proverbs sourced from various regions of Saudi Arabia. We evaluate several state-of-the-art LLMs, including multilingual and Arabic-specific models. We also provide detailed insights into their capabilities and limitations. Our results reveal notable performance gaps, particularly in tasks requiring cultural inference or contextual understanding. Our findings highlight the urgent need for dialect-aware training and culturally aligned evaluation methodologies to improve LLMs performance in real-world Arabic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。