arXiv:2502.13766cs.CL2025-02被引 19

构建全球文化知识评测基准,揭示大模型对非西方文化的显著偏见。

GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking

  • 设计跨144国的多模态文化评测框架,覆盖6大区域728个文化事件。
  • 发现模型在西方文化上表现更好,且模型越大越能识别文化起源。
  • 适合关注AI公平性、跨文化理解的研究者和开发者使用。

大型视觉语言模型(LVLMs)因其出色性能和广泛应用受到关注。然而,现有研究显示其在非西方语境下的表现不佳,且受限于文化范围窄、任务单一、模型数量少等问题。为此,我们提出GIMMICK——一个面向全球包容性的多模态多任务文化知识评测基准。该基准涵盖144个国家,代表六大全球宏观区域,包含六个任务和三个新数据集,覆盖728个独特文化事件。我们评估了20个LVLMs和11个LLMs,包括5个专有模型与26个开源模型(含不同规模)。系统分析了区域文化偏差、模型规模影响、输入模态及外部地理线索的作用。结果表明:所有模型普遍存在对西方文化的强烈偏好;模型规模与性能呈强相关;多模态输入和外部地理提示均有效提升表现;模型更擅长识别有形文化(如食物),而非无形文化(如仪式);虽能识别广泛文化来源,但在细微理解上仍不足。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have recently gained attention due to their distinctive performance and broad applicability. While it has been previously shown that their efficacy in usage scenarios involving non-Western contexts falls short, existing studies are limited in scope, covering just a narrow range of cultures, focusing exclusively on a small number of cultural aspects, or evaluating a limited selection of models on a single task only. Towards globally inclusive LVLM research, we introduce GIMMICK, an extensive multimodal benchmark designed to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. GIMMICK comprises six tasks built upon three new datasets that span 728 unique cultural events or facets on which we evaluated 20 LVLMs and 11 LLMs, including five proprietary and 26 open-weight models of all sizes. We systematically examine (1) regional cultural biases, (2) the influence of model size, (3) input modalities, and (4) external cues. Our analyses reveal strong biases toward Western cultures across models and tasks and highlight strong correlations between model size and performance, as well as the effectiveness of multimodal input and external geographic cues. We further find that models have more knowledge of tangible than intangible aspects (e.g., food vs. rituals) and that they excel in recognizing broad cultural origins but struggle with a more nuanced understanding.

多模态文化偏见评测基准LVLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。