arXiv:2509.13496cs.CVcs.LG2025-09KDD

通过注意力图发现文本生成图像中的隐性偏见并有效消除

BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image Generation

  • 用交叉注意力图揭示性别/种族与职业的语义纠缠
  • 量化纠缠程度,IoU指标显示偏见隐藏在生成过程中
  • 新方法直接修改噪声空间,同时降低分布偏差和概念耦合

偏见发现对黑箱生成模型至关重要,尤其在文本到图像(TTI)模型中。现有工作多关注输出层面的人口统计分布,但未必能保证概念表征在干预后真正解耦。我们提出BiasMap,一种模型无关的框架,用于揭示稳定扩散模型中的潜在概念级表征偏见。该方法利用交叉注意力归因图,揭示性别、种族等人口属性与职业等语义概念之间的结构纠缠,深入图像生成过程中的偏见本质。通过这些概念的归因图,我们使用交并比(IoU)量化了人口属性与语义概念的空间纠缠程度,提供了一种现有公平性检测方法无法捕捉的偏见视角。此外,我们进一步利用BiasMap进行偏见缓解,采用能量引导的扩散采样,直接修改潜在噪声空间,并在去噪过程中最小化期望软交并比(SoftIoU)。实验表明,现有公平性干预虽可减少输出分布差距,但常无法解耦概念级耦合;而我们的缓解方法能在减轻概念纠缠的同时,补充分布偏差的缓解效果。

原文摘要 · Abstract (English)

Bias discovery is critical for black-box generative models, especiall text-to-image (TTI) models. Existing works predominantly focus on output-level demographic distributions, which do not necessarily guarantee concept representations to be disentangled post-mitigation. We propose BiasMap, a model-agnostic framework for uncovering latent concept-level representational biases in stable diffusion models. BiasMap leverages cross-attention attribution maps to reveal structural entanglements between demographics (e.g., gender, race) and semantics (e.g., professions), going deeper into representational bias during the image generation. Using attribution maps of these concepts, we quantify the spatial demographics-semantics concept entanglement via Intersection over Union (IoU), offering a lens into bias that remains hidden in existing fairness discovery approaches. In addition, we further utilize BiasMap for bias mitigation through energy-guided diffusion sampling that directly modifies latent noise space and minimizes the expected SoftIoU during the denoising process. Our findings show that existing fairness interventions may reduce the output distributional gap but often fail to disentangle concept-level coupling, whereas our mitigation method can mitigate concept entanglement in image generation while complementing distributional bias mitigation.

文本生成图像偏见检测扩散模型概念解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。