揭示文生图模型内部性别偏见演化机制,发现扩散过程加剧偏见
Generated Bias: Auditing Internal Bias Dynamics of Text-To-Image Generative Models
- 提出扩散偏见与偏见放大双指标,量化多阶段生成中的内部偏见
- 实验证明文生图模型会放大性别偏见,扩散阶段本身是偏见来源之一
- 发现Stable Diffusion v2比DALL-E 2更易产生性别偏见,适合安全评估研究者
文生图(TTI)扩散模型如DALL-E和Stable Diffusion能根据文本提示生成图像。然而,已有研究指出它们会延续性别刻板印象。这些模型在多个阶段处理数据,包含多个独立训练的子模型。本文提出两项新指标,用于测量此类多阶段多模态模型内部的偏见。扩散偏见用于检测并量化模型扩散阶段引入的偏见;偏见放大则衡量从文本到图像转换过程中偏见的增强程度。实验结果表明,TTI模型会放大性别偏见,扩散过程本身即是偏见来源之一,且Stable Diffusion v2比DALL-E 2更易产生性别偏见。
原文摘要 · Abstract (English)
Text-To-Image (TTI) Diffusion Models such as DALL-E and Stable Diffusion are capable of generating images from text prompts. However, they have been shown to perpetuate gender stereotypes. These models process data internally in multiple stages and employ several constituent models, often trained separately. In this paper, we propose two novel metrics to measure bias internally in these multistage multimodal models. Diffusion Bias was developed to detect and measures bias introduced by the diffusion stage of the models. Bias Amplification measures amplification of bias during the text-to-image conversion process. Our experiments reveal that TTI models amplify gender bias, the diffusion process itself contributes to bias and that Stable Diffusion v2 is more prone to gender bias than DALL-E 2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。