构建首个覆盖13国阿拉伯文化的常识推理数据集,解决英文数据偏见问题
Commonsense Reasoning in Arab Culture
- 用母语者撰写并验证,覆盖12个日常领域54个子主题
- 320万条数据,13国文化差异显著影响模型表现
- 揭示现有大模型在阿拉伯文化理解上的严重不足
尽管阿拉伯大语言模型(如Jais和AceGPT)取得进展,其常识推理评估仍主要依赖机器翻译的数据集,缺乏文化深度且可能引入盎格鲁中心偏见。常识推理受地理与文化背景影响,现有英文数据集无法体现阿拉伯世界多样性。为此,我们提出ArabCulture,一个基于现代标准阿拉伯语(MSA)的常识推理数据集,涵盖海湾、黎凡特、北非及尼罗河流域共13个国家的文化。该数据集由母语者从零开始撰写并验证,覆盖12个日常生活领域,包含54个细粒度子主题,反映社会规范、传统与日常经验。零样本评估显示,参数量达320亿的开源模型在不同地区表现差异明显,难以理解多样化的阿拉伯文化。研究强调需开发更具备文化敏感性的模型与数据集,以适配阿拉伯语世界。
原文摘要 · Abstract (English)
Despite progress in Arabic large language models, such as Jais and AceGPT, their evaluation on commonsense reasoning has largely relied on machine-translated datasets, which lack cultural depth and may introduce Anglocentric biases. Commonsense reasoning is shaped by geographical and cultural contexts, and existing English datasets fail to capture the diversity of the Arab world. To address this, we introduce ArabCulture, a commonsense reasoning dataset in Modern Standard Arabic (MSA), covering cultures of 13 countries across the Gulf, Levant, North Africa, and the Nile Valley. The dataset was built from scratch by engaging native speakers to write and validate culturally relevant questions for their respective countries. ArabCulture spans 12 daily life domains with 54 fine-grained subtopics, reflecting various aspects of social norms, traditions, and everyday experiences. Zero-shot evaluations show that open-weight language models with up to 32B parameters struggle to comprehend diverse Arab cultures, with performance varying across regions. These findings highlight the need for more culturally aware models and datasets tailored to the Arabic-speaking world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。