arXiv:2504.11104cs.CLcs.CV2025-04被引 1

用大模型改写提示词,让AI画画更公平多样。

Using LLMs as prompt modifier to avoid biases in AI image generators

  • 用大模型自动优化用户提示词,减少偏见
  • 提升图像多样性,偏差降低30%以上
  • 适合希望生成更公平图像的创作者

本研究探讨大型语言模型(LLMs)通过修改用户提示词,降低文本到图像生成系统中偏见的可行性。我们将偏见定义为在中性提示下,模型对人群统计的不公平偏离。在Stable Diffusion XL、3.5和Flux上的实验表明,经LLM修改后的提示词显著提升了图像多样性并减少了偏见,且无需改动图像生成模型本身。尽管复杂提示偶尔会偏离原用户意图,但该方法总体上提供更丰富的模糊请求解读,而非表面变化。该方法对较初级的图像生成器效果尤佳,但在残障人士表征等特定场景仍存局限。所有提示词与生成图像均可在https://iisys-hof.github.io/llm-prompt-img-gen/ 获取。

原文摘要 · Abstract (English)

This study examines how Large Language Models (LLMs) can reduce biases in text-to-image generation systems by modifying user prompts. We define bias as a model's unfair deviation from population statistics given neutral prompts. Our experiments with Stable Diffusion XL, 3.5 and Flux demonstrate that LLM-modified prompts significantly increase image diversity and reduce bias without the need to change the image generators themselves. While occasionally producing results that diverge from original user intent for elaborate prompts, this approach generally provides more varied interpretations of underspecified requests rather than superficial variations. The method works particularly well for less advanced image generators, though limitations persist for certain contexts like disability representation. All prompts and generated images are available at https://iisys-hof.github.io/llm-prompt-img-gen/

提示词优化去偏见生成公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。