构建跨文化视觉语言模型评测基准,提升AI对全球100+国家文化的理解能力
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
- 基于19,682个文化概念构建多模态评测集CultureVerse
- 在188个国家/地区数据上验证,模型对非西方文化理解显著提升
- 适用于希望实现公平、多元文化感知的AI系统研发者
视觉语言模型虽推动人机交互进步,但受限于以西方为中心的训练数据,在文化理解上存在偏差,常误读符号、手势与文物。本文构建了覆盖19,682个文化概念、188个国别/地区、15类文化主题及3种题型的大型多模态基准CultureVerse,旨在评估并提升视觉语言模型的文化认知能力。在此基础上,提出CultureVLM系列模型,在该数据集上微调后,显著改善文化理解表现。对16个模型的评估显示,西方文化理解较强,而非洲与亚洲语境下表现较弱。微调后的模型在跨文化、跨大陆和跨数据集场景中均展现良好泛化能力,且未牺牲通用视觉语言模型基准性能。研究还揭示了文化泛化与遗忘现象,为打造更公平、更具文化敏感性的多模态AI系统奠定基础。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。