arXiv:2509.16188cs.CLcs.AI2025-09被引 9

构建多维度文化评估框架,量化大模型的文化理解能力

CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs

  • 基于文化冰山理论设计3层140维分类体系
  • 自动生成跨语言文化知识库与评测数据集
  • 揭示现有模型缺乏全面文化能力

随着大语言模型(LLMs)在多元文化环境中的广泛应用,评估其文化理解能力对于确保可信且文化契合的应用至关重要。然而,现有基准普遍存在覆盖不全、难以跨文化扩展的问题,因其框架缺乏成熟文化理论指导,且高度依赖专家手动标注。为此,我们提出CultureScope,目前最全面的文化理解评估框架。受文化冰山理论启发,设计了包含3层140个维度的新型文化知识分类体系,指导任意语言与文化下文化专属知识库及对应评测数据集的自动化构建。实验表明,该方法能有效评估文化理解能力;同时揭示现有大语言模型普遍缺乏全面文化素养,仅增加多语言数据并不能必然提升文化理解。所有代码与数据文件均可在https://github.com/HoganZinger/Culture获取。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in diverse cultural environments, evaluating their cultural understanding capability has become essential for ensuring trustworthy and culturally aligned applications. However, most existing benchmarks lack comprehensiveness and are challenging to scale and adapt across different cultural contexts, because their frameworks often lack guidance from well-established cultural theories and tend to rely on expert-driven manual annotations. To address these issues, we propose CultureScope, the most comprehensive evaluation framework to date for assessing cultural understanding in LLMs. Inspired by the cultural iceberg theory, we design a novel dimensional schema for cultural knowledge classification, comprising 3 layers and 140 dimensions, which guides the automated construction of culture-specific knowledge bases and corresponding evaluation datasets for any given languages and cultures. Experimental results demonstrate that our method can effectively evaluate cultural understanding. They also reveal that existing large language models lack comprehensive cultural competence, and merely incorporating multilingual data does not necessarily enhance cultural understanding. All code and data files are available at https://github.com/HoganZinger/Culture

文化理解评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。