评测大模型对波斯文化的理解能力,发现差距显著。
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
- 构建故事驱动的波斯文化评测数据集,由母语者参与设计。
- 最佳闭源模型与普通人表现相差11.3%,开源模型差距达21.3%。
- 适合关注跨文化NLP、多语言模型评估的研究者使用。
大型语言模型主要反映西方文化,这在很大程度上源于以英语为中心的训练数据。这种不平衡带来了严峻挑战,因为这些模型在非英语语言环境中日益广泛应用,却缺乏对其文化适应性的充分评估,波斯语即是典型例子。为此,我们提出了PerCul——一个精心构建的数据集,用于评估大模型对波斯文化的敏感性。该数据集包含基于故事的多项选择题,涵盖文化相关的复杂情境。不同于现有基准,PerCul由母语波斯语标注者共同设计,确保内容真实,防止通过翻译捷径绕过文化理解。我们评估了多个前沿多语言及专用于波斯语的大模型,为未来跨文化自然语言处理评估奠定了基础。实验显示,最佳闭源模型与普通人的表现差距为11.3%,而最佳开源模型的差距则扩大至21.3%。数据集可在此处获取:https://huggingface.co/datasets/teias-ai/percul
原文摘要 · Abstract (English)
Large language models predominantly reflect Western cultures, largely due to the dominance of English-centric training data. This imbalance presents a significant challenge, as LLMs are increasingly used across diverse contexts without adequate evaluation of their cultural competence in non-English languages, including Persian. To address this gap, we introduce PerCul, a carefully constructed dataset designed to assess the sensitivity of LLMs toward Persian culture. PerCul features story-based, multiple-choice questions that capture culturally nuanced scenarios. Unlike existing benchmarks, PerCul is curated with input from native Persian annotators to ensure authenticity and to prevent the use of translation as a shortcut. We evaluate several state-of-the-art multilingual and Persian-specific LLMs, establishing a foundation for future research in cross-cultural NLP evaluation. Our experiments demonstrate a 11.3% gap between best closed source model and layperson baseline while the gap increases to 21.3% by using the best open-weight model. You can access the dataset from here: https://huggingface.co/datasets/teias-ai/percul
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。