arXiv:2507.20973cs.LGcs.CV2025-07被引 3

用稀疏自编码器识别并抑制文本生成图像中的性别刻板印象。

Model-Agnostic Gender Bias Control for Text-to-Image Generation via Sparse Autoencoder

  • 在特征空间中用稀疏自编码器定位性别相关方向。
  • 跨多模型减少性别偏见,且不损失图像质量。
  • 无需重训练,适合各类文生图模型公平性改造。

文本到图像扩散模型常表现出性别偏见,尤其在职业与性别角色关联上存在刻板印象。本文提出SAE Debias,一种轻量级、模型无关的去偏框架。该方法不依赖CLIP过滤或提示工程,直接在特征空间操作,无需重新训练或修改结构。通过预训练于性别偏见数据集的k-稀疏自编码器,识别稀疏潜在空间中的性别相关方向,捕捉职业刻板印象。为每类职业构建一个有偏方向,并在推理时抑制该方向,以引导生成更平衡的性别结果。稀疏自编码器仅需训练一次,可重复用于多种模型。在Stable Diffusion 1.4、1.5、2.1和SDXL等模型上验证,SAE Debias显著降低性别偏见,同时保持生成质量。这是首个将稀疏自编码器应用于文生图模型去偏的研究,为构建负责任的生成式AI提供了可解释、通用的工具。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion models often exhibit gender bias, particularly by generating stereotypical associations between professions and gendered subjects. This paper presents SAE Debias, a lightweight and model-agnostic framework for mitigating such bias in T2I generation. Unlike prior approaches that rely on CLIP-based filtering or prompt engineering, which often require model-specific adjustments and offer limited control, SAE Debias operates directly within the feature space without retraining or architectural modifications. By leveraging a k-sparse autoencoder pre-trained on a gender bias dataset, the method identifies gender-relevant directions within the sparse latent space, capturing professional stereotypes. Specifically, a biased direction per profession is constructed from sparse latents and suppressed during inference to steer generations toward more gender-balanced outputs. Trained only once, the sparse autoencoder provides a reusable debiasing direction, offering effective control and interpretable insight into biased subspaces. Extensive evaluations across multiple T2I models, including Stable Diffusion 1.4, 1.5, 2.1, and SDXL, demonstrate that SAE Debias substantially reduces gender bias while preserving generation quality. To the best of our knowledge, this is the first work to apply sparse autoencoders for identifying and intervening in gender bias within T2I models. These findings contribute toward building socially responsible generative AI, providing an interpretable and model-agnostic tool to support fairness in text-to-image generation.

图像生成性别偏见去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。