首个评测大模型阿拉伯与伊斯兰文化理解力的基准任务
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
- 设计双子任务,用阿拉伯语多选题评估文化认知
- 微调后模型准确率达72.15%(文化)和84.22%(伊斯兰知识)
- 参数高效微调效果最佳,数据增强依领域而异
大型语言模型(LLMs)在预训练阶段所接触的数据主要来自网络,往往偏向资源丰富的西方语言与文化,导致其对阿拉伯及伊斯兰文化理解不足,尤其在低资源话题上表现更差。为解决此问题,我们推出PalmX 2025,首个专门评测大模型在阿拉伯与伊斯兰文化方面认知能力的共享任务。该任务包含两个子任务:通用阿拉伯文化与通用伊斯兰文化,均以现代标准阿拉伯语(MSA)呈现多选题,涵盖22个阿拉伯国家的传统、饮食、历史、宗教实践及语言表达等广泛主题。共有26支队伍参加第一子任务,19支参加第二子任务,最终分别提交9份和6份有效结果。研究发现,针对任务进行微调显著提升性能;表现最佳系统在文化问题上达到72.15%准确率,在伊斯兰知识上达84.22%。参数高效微调成为主流且最有效的方法,而数据增强的效果则依赖具体领域。
原文摘要 · Abstract (English)
Large Language Models (LLMs) inherently reflect the vast data distributions they encounter during their pre-training phase. As this data is predominantly sourced from the web, there is a high chance it will be skewed towards high-resourced languages and cultures, such as those of the West. Consequently, LLMs often exhibit a diminished understanding of certain communities, a gap that is particularly evident in their knowledge of Arabic and Islamic cultures. This issue becomes even more pronounced with increasingly under-represented topics. To address this critical challenge, we introduce PalmX 2025, the first shared task designed to benchmark the cultural competence of LLMs in these specific domains. The task is composed of two subtasks featuring multiple-choice questions (MCQs) in Modern Standard Arabic (MSA): General Arabic Culture and General Islamic Culture. These subtasks cover a wide range of topics, including traditions, food, history, religious practices, and language expressions from across 22 Arab countries. The initiative drew considerable interest, with 26 teams registering for Subtask 1 and 19 for Subtask 2, culminating in nine and six valid submissions, respectively. Our findings reveal that task-specific fine-tuning substantially boosts performance over baseline models. The top-performing systems achieved an accuracy of 72.15% on cultural questions and 84.22% on Islamic knowledge. Parameter-efficient fine-tuning emerged as the predominant and most effective approach among participants, while the utility of data augmentation was found to be domain-dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。