arXiv:2512.03173cs.CYcs.AI2025-12AAAI被引 2

用功能映射重构物品多样性,让视觉语言模型更公平。

Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping

  • 按物品功能而非外观分类,跨文化统一标注
  • 低收入群体性能差距缩小6个百分点,显著提升公平性
  • 适合关注AI公平性与跨文化数据构建的研究者

文化塑造人们使用的物品及其用途,但主流视觉-语言(VL)数据集普遍存在文化偏见,过度偏向高收入、西方背景。这种不平衡降低了模型泛化能力,加剧了性能差异,尤其影响低收入和非西方群体。为此,我们提出一种以功能为核心的全新框架,通过功能维度对物品进行跨文化、跨经济背景的分类。我们基于Dollar Street数据集构建了文化可用性图谱(Culture Affordance Atlas),涵盖46种功能和288个物品,公开于https://lit.eecs.umich.edu/CultureAffordance-Atlas/index.html。利用CLIP模型进行大量实证分析表明,以功能为中心的标签可将高低收入群体间的性能差距中位数降低6个百分点(统计显著),显著提升对低收入情境的模型有效性。此外,我们的分析揭示了许多在主流VL数据集中常被忽略的文化关键物品。本工作为构建包容性视觉-语言数据集和公平的人工智能系统提供了可扩展路径。

原文摘要 · Abstract (English)

Culture shapes the objects people use and for what purposes, yet mainstream Vision-Language (VL) datasets frequently exhibit cultural biases, disproportionately favoring higher-income, Western contexts. This imbalance reduces model generalizability and perpetuates performance disparities, especially impacting lower-income and non-Western communities. To address these disparities, we propose a novel function-centric framework that categorizes objects by the functions they fulfill, across diverse cultural and economic contexts. We implement this framework by creating the Culture Affordance Atlas, a re-annotated and culturally grounded restructuring of the Dollar Street dataset spanning 46 functions and 288 objects publicly available at https://lit.eecs.umich.edu/CultureAffordance-Atlas/index.html. Through extensive empirical analyses using the CLIP model, we demonstrate that function-centric labels substantially reduce socioeconomic performance gaps between high- and low-income groups by a median of 6 pp (statistically significant), improving model effectiveness for lower-income contexts. Furthermore, our analyses reveals numerous culturally essential objects that are frequently overlooked in prominent VL datasets. Our contributions offer a scalable pathway toward building inclusive VL datasets and equitable AI systems.

视觉语言公平性文化多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。