通过优化参数与提示结构,显著降低文生图模型的偏见和能耗。
SustainDiffusion: Optimising the Social and Environmental Sustainability of Stable Diffusion Models
- 基于搜索算法寻找最优超参数与提示组合
- 性别偏见降68%,种族偏见降59%,能耗减48%
- 无需修改模型架构,适合注重伦理与环保的开发者
文本到图像生成模型应用广泛,其中开源的Stable Diffusion(SD)每年生成超过120亿张图像。然而其广泛应用引发社会与环境可持续性担忧。本文提出SustainDiffusion,一种基于搜索的方法,旨在降低生成图像中的性别与种族偏见,同时减少图像生成所需的能源消耗。该方法在不改变模型架构的前提下,通过优化超参数与提示结构,实现与原始SD模型相当的图像质量。我们在6个基线模型上使用56个不同提示进行了全面评估,结果表明,SustainDiffusion可使SD3的性别偏见降低68%,种族偏见降低59%,能源消耗(CPU+GPU总和)减少48%。且结果在多次运行中保持一致,并可推广至多种提示。本研究证明,在不进行微调或架构更改的情况下,提升文本到图像模型的社会与环境可持续性是可行的。
原文摘要 · Abstract (English)
Background: Text-to-image generation models are widely used across numerous domains. Among these models, Stable Diffusion (SD) - an open-source text-to-image generation model - has become the most popular, producing over 12 billion images annually. However, the widespread use of these models raises concerns regarding their social and environmental sustainability. Aims: To reduce the harm that SD models may have on society and the environment, we introduce SustainDiffusion, a search-based approach designed to enhance the social and environmental sustainability of SD models. Method: SustainDiffusion searches the optimal combination of hyperparameters and prompt structures that can reduce gender and ethnic bias in generated images while also lowering the energy consumption required for image generation. Importantly, SustainDiffusion maintains image quality comparable to that of the original SD model. Results: We conduct a comprehensive empirical evaluation of SustainDiffusion, testing it against six different baselines using 56 different prompts. Our results demonstrate that SustainDiffusion can reduce gender bias in SD3 by 68%, ethnic bias by 59%, and energy consumption (calculated as the sum of CPU and GPU energy) by 48%. Additionally, the outcomes produced by SustainDiffusion are consistent across multiple runs and can be generalised to various prompts. Conclusions: With SustainDiffusion, we demonstrate how enhancing the social and environmental sustainability of text-to-image generation models is possible without fine-tuning or changing the model's architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。