构建涵盖22个阿拉伯国家的多语言对话数据集,提升大模型对阿拉伯文化的理解能力。
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
- 覆盖22个阿拉伯国家,包含标准语与方言的指令-响应对
- 评估多个前沿大模型在文化与方言上的表现,发现代表性不均
- 开源标注规范与代码,支持复现与后续研究
随着大语言模型日益融入日常生活,确保其文化敏感性与包容性至关重要。我们推出一个历时一年、由阿拉伯世界44位研究人员共同参与的社区驱动项目,覆盖全部22个阿拉伯国家。数据集包含现代标准阿拉伯语(MSA)和方言(DA)的指令-响应对,涵盖20个多样化主题。通过该数据集评估多个前沿大模型的文化与方言能力,发现闭源模型整体表现良好但仍存缺陷,小规模开源模型面临更大挑战;部分国家(如埃及、阿联酋)代表性较强,而伊拉克、毛里塔尼亚、也门等则相对不足。所有标注指南、代码与数据均公开可获取,支持复现与持续研究。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce our dataset, a year-long community-driven project covering all 22 Arab countries. The dataset includes instructions (input, response pairs) in both Modern Standard Arabic (MSA) and dialectal Arabic (DA), spanning 20 diverse topics. Built by a team of 44 researchers across the Arab world, all of whom are authors of this paper, our dataset offers a broad, inclusive perspective. We use our dataset to evaluate the cultural and dialectal capabilities of several frontier LLMs, revealing notable limitations. For instance, while closed-source LLMs generally exhibit strong performance, they are not without flaws, and smaller open-source models face greater challenges. Moreover, certain countries (e.g., Egypt, the UAE) appear better represented than others (e.g., Iraq, Mauritania, Yemen). Our annotation guidelines, code, and data for reproducibility are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。