arXiv:2510.05931cs.CLcs.CY2025-10Conference of the …被引 13

用人类学视角重构语言模型文化评估标准

Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens

  • 提出四类文化基准框架,分析其对文化的理解偏差
  • 发现20个基准普遍存在国家等同文化等六大问题
  • 建议引入真实叙事与社区参与,实现情境化评估

大语言模型的文化评估日益重要,但现有基准常将文化简化为静态事实或同质价值,与人类学强调文化动态性、历史性及实践性的观点相悖。本文提出一个四部分框架,用于分类基准如何界定文化(知识、偏好、表现、偏见)。基于该框架,我们对20个文化基准进行定性分析,识别出六项常见方法论缺陷:将国家等同于文化、忽视文化内部多样性、依赖简化问卷等。借鉴人类学方法,提出三项改进策略:融入真实生活叙事与场景、让文化群体参与设计与验证、在具体语境中评估模型表现。目标是推动文化基准超越静态记忆任务,更准确反映模型应对复杂文化情境的能力。

原文摘要 · Abstract (English)

Cultural evaluation of large language models has become increasingly important, yet current benchmarks often reduce culture to static facts or homogeneous values. This view conflicts with anthropological accounts that emphasize culture as dynamic, historically situated, and enacted in practice. To analyze this gap, we introduce a four-part framework that categorizes how benchmarks frame culture, such as knowledge, preference, performance, or bias. Using this lens, we qualitatively examine 20 cultural benchmarks and identify six recurring methodological issues, including treating countries as cultures, overlooking within-culture diversity, and relying on oversimplified survey formats. Drawing on established anthropological methods, we propose concrete improvements: incorporating real-world narratives and scenarios, involving cultural communities in design and validation, and evaluating models in context rather than isolation. Our aim is to guide the development of cultural benchmarks that go beyond static recall tasks and more accurately capture the responses of the models to complex cultural situations.

文化评估人类学LLM基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。