构建阿拉伯语方言对话数据集,评估大模型在文化语境下的表现差异
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

- 构建覆盖13国的阿拉伯语对话数据集,含标准语与方言双版本
- 模型在方言任务上表现显著低于标准语,三类任务均存在性能差距
- 适合研究多语言文化推理、本地化大模型的学者与开发者参考
当前大模型在文化推理方面的评估仍存在明显缺口,现有阿拉伯语基准多聚焦于现代标准阿拉伯语(MSA)的短文本片段,忽视了对话中自然涌现的文化细节。为此,我们提出 ArabCulture-Dialogue,一个以文化为根基的对话数据集,涵盖13个阿拉伯语国家,包含每国方言及对应的标准语,覆盖12个日常生活话题和54个细粒度子话题。基于该数据集,我们设计三项评测任务:(i) 多选文化推理,(ii) MSA与方言间的机器翻译,(iii) 方言引导生成。实验表明,模型在方言设置下的表现普遍低于标准语设置,三类任务均存在显著性能差距。
原文摘要 · Abstract (English)
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. We utilize the dataset to form three benchmarking tasks: (i) multiple-choice cultural reasoning, (ii) machine translation between MSA and dialects, and (iii) dialect-steering generation. Our experiments indicate that the performance gap between MSA and Arabic dialects still exists, whereby the models perform worse on all three tasks in the dialectal setup, compared to the MSA one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。