提出五维文化维度框架,评估视觉语言模型的文化理解能力
Evaluation of Cultural Competence of Vision-Language Models
- 基于视觉文化研究方法构建五维文化分析框架
- 系统性识别图像中的文化细微差异,弥补现有评估短板
- 适合研究AI文化偏见或跨文化AI应用的学者与开发者
现代视觉语言模型在文化能力评测中表现不佳。随着基于VLM的应用日益多样化,亟需深入理解其如何编码文化细微差别。尽管部分问题已受关注,但尚缺乏系统性的框架来识别和标注图像中的文化维度。本文主张借鉴视觉文化研究(文化研究、符号学、视觉研究)的基础方法,提出五个对应文化维度的分析框架,以实现对VLM文化能力更全面的评估。
原文摘要 · Abstract (English)
Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural nuances. While individual aspects of this problem have been studied, we still lack a comprehensive framework for systematically identifying and annotating the nuanced cultural dimensions present in images for VLMs. This position paper argues that foundational methodologies from visual culture studies (cultural studies, semiotics, and visual studies) are necessary for cultural analysis of images. Building upon this review, we propose a set of five frameworks, corresponding to cultural dimensions, that must be considered for a more complete analysis of the cultural competencies of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。