arXiv:2503.17485cs.CLcs.AI2025-03被引 20

构建沙特文化评测基准,检验大模型在本地文化理解上的短板。

SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia

  • 设计覆盖5个地区的文化问答数据集,含多类型题目
  • 所有模型在区域特异性问题上表现显著下降
  • 适合评估大模型在本土文化场景下的适应能力

大型语言模型(LLMs)在自然语言处理中表现出色,但往往难以准确捕捉和反映文化细节。本研究聚焦沙特阿拉伯,一个方言多样、文化丰富的国家,提出SaudiCulture——一个用于评估大模型在沙特特定地理与文化背景下文化胜任力的新基准。该数据集涵盖西、东、南、北、中心五大地区及全境通用的问题,覆盖饮食、服饰、娱乐、节日、手工艺等广泛文化领域。题目形式包括开放题、单选、多选,部分需多个正确答案,区分通用文化知识与区域专有内容。对GPT-4、Llama 3.3、FANAR、Jais、AceGPT五款模型的评估显示,面对高度专业化或地域性问题时,所有模型性能均明显下降,尤其在多正确答案题目上表现不佳。部分文化类别更易识别,暴露出模型文化理解的不一致性。结果强调了在训练中融入区域知识对提升模型文化胜任力的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing; however, they often struggle to accurately capture and reflect cultural nuances. This research addresses this challenge by focusing on Saudi Arabia, a country characterized by diverse dialects and rich cultural traditions. We introduce SaudiCulture, a novel benchmark designed to evaluate the cultural competence of LLMs within the distinct geographical and cultural contexts of Saudi Arabia. SaudiCulture is a comprehensive dataset of questions covering five major geographical regions, such as West, East, South, North, and Center, along with general questions applicable across all regions. The dataset encompasses a broad spectrum of cultural domains, including food, clothing, entertainment, celebrations, and crafts. To ensure a rigorous evaluation, SaudiCulture includes questions of varying complexity, such as open-ended, single-choice, and multiple-choice formats, with some requiring multiple correct answers. Additionally, the dataset distinguishes between common cultural knowledge and specialized regional aspects. We conduct extensive evaluations on five LLMs, such as GPT-4, Llama 3.3, FANAR, Jais, and AceGPT, analyzing their performance across different question types and cultural contexts. Our findings reveal that all models experience significant performance declines when faced with highly specialized or region-specific questions, particularly those requiring multiple correct responses. Additionally, certain cultural categories are more easily identifiable than others, further highlighting inconsistencies in LLMs cultural understanding. These results emphasize the importance of incorporating region-specific knowledge into LLMs training to enhance their cultural competence.

文化理解大模型评测沙特研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。