arXiv:2410.15453cs.CLcs.CV2024-10NAACL被引 12

测试视觉语言模型对文化概念的适应能力,发现其理解力有限。

CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts

  • 构建新基准CROPE,区分模型训练知识与推理时上下文知识。
  • 模型在文化特定概念上表现远差于通用概念,参数化知识不足。
  • 模型难以融合多模态信息绑定文化概念,适合研究跨文化AI的学者。

随着视觉语言模型(VLMs)在全球范围普及,评估其文化理解能力成为关键挑战。本文提出CROPE,一个用于探测文化特定概念知识并评估上下文适应能力的视觉问答基准。该基准可区分模型在训练中习得的参数化知识与推理时通过图文描述提供的上下文知识。对多个先进开源VLMs的评估显示,在参数化设置下,模型在文化特定概念上的表现显著低于通用概念。此外,上下文实验表明,模型难以有效利用多模态信息,且无法将文化特定概念与其图像表征正确关联。研究揭示了当前VLMs在文化理解和适应性方面的局限,亟需改进以实现更包容的文化智能。

原文摘要 · Abstract (English)

As Vision and Language models (VLMs) are reaching users across the globe, assessing their cultural understanding has become a critical challenge. In this paper, we introduce CROPE, a visual question answering benchmark designed to probe the knowledge of culture-specific concepts and evaluate the capacity for cultural adaptation through contextual information. This allows us to distinguish between parametric knowledge acquired during training and contextual knowledge provided during inference via visual and textual descriptions. Our evaluation of several state-of-the-art open VLMs shows large performance disparities between culture-specific and common concepts in the parametric setting. Moreover, experiments with contextual knowledge indicate that models struggle to effectively utilize multimodal information and bind culture-specific concepts to their depictions. Our findings reveal limitations in the cultural understanding and adaptability of current VLMs that need to be addressed toward more culturally inclusive models.

视觉语言模型文化理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。