arXiv:2608.29847cs.CVcs.CL2026-08中稿 · EMNLP

测试文生图模型在不同语境下职业刻板印象是否持续存在

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

论文配图:ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
图 1 · 摘自论文原文
  • 构建可控语境评估框架,分离语义上下文影响
  • 跨角色属性集中度提升0.047,刻板印象在无关语境仍顽固
  • 适合关注生成模型偏见与评估方法的研究者

文生图模型会学习概念间关联(如职业与视觉特征),这些关联可能引发刻板偏见。本文提出ContextBias评估框架和ContextBench基准,涵盖92种职业与1,656个语义控制提示,旨在隔离上下文变化对角色相关视觉表征的影响。在66,240张生成图像上评估四款顶尖模型发现:将职业置于语义无关语境中,并未削弱角色关联属性,反而导致跨角色属性集中度增加(合并偏差指数 +0.047)。性别、服装特征及职业工具在无上下文、相关与无关条件下均高度普遍存在,且对提示重构具有鲁棒性。场景构图与镜头角度表现出最强的上下文敏感性。结果揭示一种在无上下文评估中难以察觉的刻板偏见持续现象,强调需在偏见评测中引入受控的上下文变异。代码与数据集:https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias

原文摘要 · Abstract (English)

Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open question in this area is whether these associations are stable or change when visual representations of people in professional roles are placed in different prompted contexts. We introduce ContextBias, a controlled evaluation framework, and ContextBench, a benchmark spanning 92 roles and 1,656 semantically controlled prompts, designed to isolate the effect of contextual variation on role-linked visual representations. Evaluating four state-of-the-art models on 66,240 generated images, we find that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$). Demographic cues, characteristic garments, and role-specific tools remain highly prevalent across context-free, related, and unrelated conditions, and are robust to semantic prompt reformulation. Scene composition and camera framing show the greatest context-sensitivity. These findings reveal a form of stereotypical persistence that remains largely invisible to context-free evaluations, highlighting the need for controlled contextual variation in bias benchmarking. Code and dataset: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias

文生图偏见评估刻板印象上下文敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。