arXiv:2605.31556cs.CVcs.AI2026-05中稿 · EMNLP

视觉语言模型在模糊图像中会隐性偏袒男性,即使内部编码了女性关联。

Vision-Language Models Suppress Female Representations Under Ambiguous Input

论文配图:Vision-Language Models Suppress Female Representations Under Ambiguous Input
图 1 · 摘自论文原文
  • 用零样本度量法探测模型内部语义关联,发现性别隐含偏见
  • 800张模糊图像测试中,模型输出男性占比超80%,但内部仍存女性关联
  • 适合关注模型公平性与视觉表征机制的研究者

对齐训练使视觉语言模型(VLMs)能避免表达人口统计学偏见,在性别明确时表现良好。然而,对于模糊输入(如全装束的工人、背面身影)——这类情况在实际中常见却少被研究——则知之甚少。我们发现,轻微提示压力下,模型在模糊图像上会暴露职业-性别默认倾向,甚至在强烈女性化的职业中也倾向于输出男性。这些输出是否反映模型内部真实编码?我们引入LALS(潜在关联学习得分),一种零样本度量方法,将视觉标记激活投影至文本嵌入空间,以逐层测量概念关联。在15个职业、超过800张性别模糊图像和四种VLMs上,内部表征与输出常出现系统性解耦:模型内部常编码女性关联,但输出却为男性。分层分析显示,男性能全程增强,而女性信号在中间层达峰后被抑制。颜色消融实验表明,服装色彩等文化负载视觉线索进一步调节内部关联。

原文摘要 · Abstract (English)

Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in full gear, a figure seen from behind), cases common in practice yet rarely studied. We find that minimal prompting pressure exposes occupation-gender defaults when prompting ambiguous input images, with models collapsing to male even for strongly female-stereotyped occupations. But do these outputs reflect what models actually encode internally? We introduce LALS (Latent Association Leaning Score), a zero-shot metric that projects visual-token activations into the model's text-embedding space to measure concept associations per token and layer. Across 15 occupations, over 800 gender-ambiguous images, and four VLMs, internal representations and outputs often become systematically decoupled: models often encode a female association internally yet output male. Layer-wise analysis reveals an asymmetric filter: male signal amplifies end-to-end while female signal peaks mid-network and is suppressed before generation. A color ablation shows that culturally loaded visual cues such as clothing color further modulate these internal associations.

模型偏见视觉语言模型公平性内部表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。