自动搜索提示词暴露文生图模型隐藏偏见。
Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search
- 用语言模型与属性分类器协同生成能放大偏见的提示词。
- 在Stable Diffusion 1.5和去偏模型中发现多种未记录的隐蔽偏见。
- 生成提示可被普通用户直接使用,适合评估模型公平性。
文生图扩散模型虽视觉质量优异,但在性别、种族、年龄等敏感属性上屡现社会偏见。现有缓解方法依赖人工或大模型生成的提示数据集,但存在成本高且易遗漏潜在偏见提示的问题。本文提出偏见引导提示搜索(BGPS)框架,通过指令化语言模型生成属性中立提示,并利用图像内部表征的属性分类器引导解码过程,聚焦于加剧特定属性的提示空间。在Stable Diffusion 1.5及一个先进去偏模型上进行广泛实验,发现一系列此前未记录的细微偏见,严重恶化公平性指标。关键在于,所发现提示具有可解释性,可由普通用户直接输入,且在困惑度上优于主流硬提示优化方法。该研究揭示文生图模型漏洞,而BGPS扩展了偏见搜索范围,可作为新型偏见评估工具。
原文摘要 · Abstract (English)
Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been repeatedly shown to exhibit social biases across sensitive attributes such as gender, race and age. To mitigate these biases, existing approaches frequently depend on curated prompt datasets - either manually constructed or generated with large language models (LLMs) - as part of their training and/or evaluation procedures. Beside the curation cost, this also risks overlooking unanticipated, less obvious prompts that trigger biased generation, even in models that have undergone debiasing. In this work, we introduce Bias-Guided Prompt Search (BGPS), a framework that automatically generates prompts that aim to maximize the presence of biases in the resulting images. BGPS comprises two components: (1) an LLM instructed to produce attribute-neutral prompts and (2) attribute classifiers acting on the TTI's internal representations that steer the decoding process of the LLM toward regions of the prompt space that amplify the image attributes of interest. We conduct extensive experiments on Stable Diffusion 1.5 and a state-of-the-art debiased model and discover an array of subtle and previously undocumented biases that severely deteriorate fairness metrics. Crucially, the discovered prompts are interpretable, i.e they may be entered by a typical user, quantitatively improving the perplexity metric compared to a prominent hard prompt optimization counterpart. Our findings uncover TTI vulnerabilities, while BGPS expands the bias search space and can act as a new evaluation tool for bias mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。