arXiv:2602.12659cs.CVcs.AI2026-02

构建首个覆盖印度全境的平衡人脸数据集,用于检测和缓解视觉语言模型的地域偏见。

IndicFairFace: Balanced Indian Face Dataset for Auditing and Mitigating Geographical Bias in Vision-Language Models

  • 基于14,400张印度各地人脸图像,按州和性别均匀分布。
  • 发现主流CLIP模型存在显著的国内地理偏见,经去偏后检索准确率下降不足1.5%。
  • 为研究印度语境下视觉语言模型的地域偏差提供首个基准数据集。

视觉语言模型(VLMs)常从大规模网络数据中继承并放大社会偏见,其中印度群体尤其被简化代表。现有公平性数据集虽在种族与性别上实现平衡,却将印度视为单一整体,忽视其28个邦和8个中央直辖区的内部多样性,导致表征与地理偏见。为此,我们提出IndicFairFace,一个包含14,400张图像的新平衡人脸数据集,图像来源为Wikimedia Commons及开源网络资源,按州与性别均匀采样。利用该数据集,我们量化了主流基于CLIP的VLM中的国内地理偏见,并通过后处理迭代零空间投影方法有效降低偏见。结果显示,该去偏方法对原有嵌入空间影响极小,基准数据集上的平均检索准确率下降低于1.5%。本工作确立IndicFairFace为首个用于研究印度语境下视觉语言模型地理偏见的基准。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are known to inherit and amplify societal biases from their web-scale training data with Indian being particularly misrepresented. Existing fairness-aware datasets have significantly improved demographic balance across global race and gender groups, yet they continue to treat Indian as a single monolithic category. The oversimplification ignores the vast intra-national diversity across 28 states and 8 Union Territories of India and leads to representational and geographical bias. To address the limitation, we present IndicFairFace, a novel and balanced face dataset comprising 14,400 images representing geographical diversity of India. Images were sourced ethically from Wikimedia Commons and open-license web repositories and uniformly balanced across states and gender. Using IndicFairFace, we quantify intra-national geographical bias in prominent CLIP-based VLMs and reduce it using post-hoc Iterative Nullspace Projection debiasing approach. We also show that the adopted debiasing approach does not adversely impact the existing embedding space as the average drop in retrieval accuracy on benchmark datasets is less than 1.5 percent. Our work establishes IndicFairFace as the first benchmark to study geographical bias in VLMs for the Indian context.

视觉语言模型地理偏见数据集去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。