发现视觉语言模型的内在偏见会传递到零样本检索任务中,影响公平性。
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
- 通过内在表征与外在检索结果的相关性,系统分析偏见传播路径。
- 平均相关系数ρ=0.83±0.10,跨114次实验、6类社会群体均显著。
- 大模型偏见传播更强,低代表群体结果更不稳定,适合关注公平性的研究者。
为构建公平的AI系统,需理解基础编码型视觉语言模型(VLMs)中固有的社会群体偏见如何在下游任务中体现。本研究证明,VLM表征中的内在偏见会系统性“传递”至零样本检索任务,揭示偏见如何深层影响模型输出。我们提出一个受控框架,通过关联(a)表征空间内的内在偏见度量与(b)零样本文本到图像(TTI)和图像到文本(ITT)检索的外在偏见度量,衡量偏见传播。结果表明内在与外在偏见间存在显著相关性,平均ρ = 0.83 ± 0.10。该模式在114次分析中保持一致,涵盖两个检索方向、六类社会群体及三种不同VLMs。值得注意的是,更大/性能更好的模型表现出更强的偏见传播,引发对日益复杂模型的担忧。我们的框架引入基准评估任务,用于测量群体与价值信号的传播。调查发现,代表性不足群体的传播更弱,进一步加剧其模型结果的偏差。
原文摘要 · Abstract (English)
To build fair AI systems we need to understand how social-group biases intrinsic to foundational encoder-based vision-language models (VLMs) manifest in biases in downstream tasks. In this study, we demonstrate that intrinsic biases in VLM representations systematically ``carry over'' or propagate into zero-shot retrieval tasks, revealing how deeply rooted biases shape a model's outputs. We introduce a controlled framework to measure this propagation by correlating (a) intrinsic measures of bias in the representational space with (b) extrinsic measures of bias in zero-shot text-to-image (TTI) and image-to-text (ITT) retrieval. Results show substantial correlations between intrinsic and extrinsic bias, with an average $ρ$ = 0.83 $\pm$ 0.10. This pattern is consistent across 114 analyses, both retrieval directions, six social groups, and three distinct VLMs. Notably, we find that larger/better-performing models exhibit greater bias propagation, a finding that raises concerns given the trend towards increasingly complex AI models. Our framework introduces baseline evaluation tasks to measure the propagation of group and valence signals. Investigations reveal that underrepresented groups experience less robust propagation, further skewing their model-related outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。