探究视觉语言模型在跨国文化中的偏见,揭示其隐含的种族与身份刻板印象
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
- 设计三类检索任务,评估模型对不同国家与种族、特质、外貌的关联理解
- 发现模型普遍存在将特定种族或外貌与国家绑定的系统性偏见
- 适用于关注AI公平性、跨文化AI伦理的研究者与开发者
视觉语言模型(VLMs)在多样文化环境中应用日益广泛,但其内部偏见尚不明确。本文提出一种新框架,系统评估VLMs在不同国家间对种族、性别及身体特征的文化差异与偏见编码。引入三个基于检索的任务:(1) 种族到国家检索,考察东亚、白人、中东、拉美、南亚及黑人等群体与国家间的关联;(2) 个人特质到国家检索,通过图像与特质提示(如聪明、诚实、罪犯、暴力)分析潜在刻板印象;(3) 身体特征到国家检索,聚焦瘦、年轻、肥胖、年老等视觉属性与国家的文化关联。结果揭示模型中持续存在的偏见,表明视觉表征可能无意中强化社会刻板印象。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly deployed in diverse cultural contexts, yet their internal biases remain poorly understood. In this work, we propose a novel framework to systematically evaluate how VLMs encode cultural differences and biases related to race, gender, and physical traits across countries. We introduce three retrieval-based tasks: (1) Race to Country retrieval, which examines the association between individuals from specific racial groups (East Asian, White, Middle Eastern, Latino, South Asian, and Black) and different countries; (2) Personal Traits to Country retrieval, where images are paired with trait-based prompts (e.g., Smart, Honest, Criminal, Violent) to investigate potential stereotypical associations; and (3) Physical Characteristics to Country retrieval, focusing on visual attributes like skinny, young, obese, and old to explore how physical appearances are culturally linked to nations. Our findings reveal persistent biases in VLMs, highlighting how visual representations may inadvertently reinforce societal stereotypes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。