arXiv:2505.14729cs.CV2025-05被引 3

检测视觉语言模型在识别国家图像时的文化偏见差异

Uncovering Cultural Representation Disparities in Vision-Language Models

  • 用多国图像数据集测试模型跨文化识别能力
  • 不同国家准确率差异显著,最高差达40%
  • 适合关注AI公平性与数据偏见的研究者

视觉语言模型(VLMs)在多种任务中表现优异,但其潜在偏见引发关注。本文通过国家层面的图像识别任务,评估主流VLMs在Country211数据集上的表现。采用开放式提问和多选题(包括多语言、对抗性等挑战设置),分析不同国家与题型下的准确率差异。结果表明,模型性能存在显著不均衡,最高国家间差距达40%,说明尽管模型具备较强视觉理解能力,仍继承预训练数据分布带来的文化偏见,影响其在全球多样化场景下的泛化能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated impressive capabilities across a range of tasks, yet concerns about their potential biases exist. This work investigates the extent to which prominent VLMs exhibit cultural biases by evaluating their performance on an image-based country identification task at a country level. Utilizing the geographically diverse Country211 dataset, we probe several large vision language models (VLMs) under various prompting strategies: open-ended questions, multiple-choice questions (MCQs) including challenging setups like multilingual and adversarial settings. Our analysis aims to uncover disparities in model accuracy across different countries and question formats, providing insights into how training data distribution and evaluation methodologies might influence cultural biases in VLMs. The findings highlight significant variations in performance, suggesting that while VLMs possess considerable visual understanding, they inherit biases from their pre-training data and scale that impact their ability to generalize uniformly across diverse global contexts.

视觉语言模型文化偏见公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。