用帕累托前沿评估文生图模型的公平性与质量权衡,可找出更优参数配置。
A Framework for Benchmarking Fairness-Utility Trade-offs in Text-to-Image Models via Pareto Frontiers
- 通过帕累托前沿分析不同去偏方法的超参数,系统评估公平性与图像质量
- 多数默认参数在公平性-质量空间中被优于,存在更优配置
- 适用于评估和优化文生图模型的去偏策略,适合负责任AI研究者
文生图模型的公平性需在消除社会偏见的同时保持视觉质量,这对负责任的人工智能至关重要。当前评估方法依赖主观判断或有限对比,难以全面、可复现地评估公平性与实用性。现有方法多采用人工视觉检查,易出错且难重复。本文提出基于帕累托前沿的评估框架,通过超参数化去偏方法,在公平性与实用性之间建立可比较的权衡关系。我们使用归一化香农熵衡量公平性,ClipScore衡量实用性,评估了Stable Diffusion、Fair Diffusion、SDXL、DeCoDi和FLUX等模型。结果表明,多数默认超参数配置在公平性-实用性空间中为被支配解,可轻松找到更优参数组合。
原文摘要 · Abstract (English)
Achieving fairness in text-to-image generation demands mitigating social biases without compromising visual fidelity, a challenge critical to responsible AI. Current fairness evaluation procedures for text-to-image models rely on qualitative judgment or narrow comparisons, which limit the capacity to assess both fairness and utility in these models and prevent reproducible assessment of debiasing methods. Existing approaches typically employ ad-hoc, human-centered visual inspections that are both error-prone and difficult to replicate. We propose a method for evaluating fairness and utility in text-to-image models using Pareto-optimal frontiers across hyperparametrization of debiasing methods. Our method allows for comparison between distinct text-to-image models, outlining all configurations that optimize fairness for a given utility and vice-versa. To illustrate our evaluation method, we use Normalized Shannon Entropy and ClipScore for fairness and utility evaluation, respectively. We assess fairness and utility in Stable Diffusion, Fair Diffusion, SDXL, DeCoDi, and FLUX text-to-image models. Our method shows that most default hyperparameterizations of the text-to-image model are dominated solutions in the fairness-utility space, and it is straightforward to find better hyperparameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。