揭示文本生成图像模型中偏见的相互作用,避免修复一个偏见时加剧另一个。
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
- 基于反事实框架构建偏见交互的因果图,量化多维度偏见关系。
- 发现修改某一偏见维度时,其他维度偏差变化与理想分布相关性达0.69。
- 适用于评估模型公平性、选择最优去偏方向,尤其适合研究交叉偏见者。
文本到图像(TTI)模型中的偏见常被视作独立存在,但实际可能深度交织。调整某一维度(如种族或年龄)的偏见,可能意外影响另一维度(如性别),导致偏差缓解或恶化。理解这种相互依赖对设计更公平的生成模型至关重要,但量化此类效应仍具挑战。本文提出BiasConnect,一种新工具,用于分析和量化TTI模型中的偏见交互。方法基于反事实框架,生成针对特定文本提示的成对因果图,揭示偏见交互的内在结构。同时提供实证估计,显示当某一偏见被调整时,其他偏见维度向或远离理想分布的移动趋势。该估计与偏见缓解后的相互依赖观察高度相关(+0.69)。我们展示了BiasConnect在选择最优去偏轴、比较不同TTI模型所学习的依赖关系,以及理解交叉社会偏见在模型中被放大的作用。
原文摘要 · Abstract (English)
The biases exhibited by Text-to-Image (TTI) models are often treated as if they are independent, but in reality, they may be deeply interrelated. Addressing bias along one dimension, such as ethnicity or age, can inadvertently influence another dimension, like gender, either mitigating or exacerbating existing disparities. Understanding these interdependencies is crucial for designing fairer generative models, yet measuring such effects quantitatively remains a challenge. In this paper, we aim to address these questions by introducing BiasConnect, a novel tool designed to analyze and quantify bias interactions in TTI models. Our approach leverages a counterfactual-based framework to generate pairwise causal graphs that reveals the underlying structure of bias interactions for the given text prompt. Additionally, our method provides empirical estimates that indicate how other bias dimensions shift toward or away from an ideal distribution when a given bias is modified. Our estimates have a strong correlation (+0.69) with the interdependency observations post bias mitigation. We demonstrate the utility of BiasConnect for selecting optimal bias mitigation axes, comparing different TTI models on the dependencies they learn, and understanding the amplification of intersectional societal biases in TTI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。