评估文生图模型的公平性、多样性与可靠性,发现其易受提示词干扰。
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
- 通过扰动嵌入空间,检测模型在不同输入下的不可靠行为。
- 发现移除提示词会显著影响生成控制,暴露模型偏见。
- 适合关注AI伦理、模型安全与可信生成的研究者使用。
多模态生成模型的快速普及引发了对其可靠性、公平性及滥用风险的广泛关注。尽管文生图模型能生成高质量、用户引导的内容,但常表现出不可预测的行为和脆弱性,可能被用于操纵类别或概念表征。为此,我们提出一种评估框架,通过分析嵌入空间中的全局与局部扰动,识别触发不可靠或偏见行为的输入。该方法深入评估生成多样性(衡量学习概念的视觉表示广度)与生成公平性(在低引导设置下,移除提示词对生成控制的影响)。此外,该方法为检测注入偏见的模型及追踪偏见来源提供基础。代码已公开于 https://github.com/JJ-Vice/T2I_Fairness_Diversity_Reliability。
原文摘要 · Abstract (English)
The rapid proliferation of multimodal generative models has sparked critical discussions on their reliability, fairness and potential for misuse. While text-to-image models excel at producing high-fidelity, user-guided content, they often exhibit unpredictable behaviors and vulnerabilities that can be exploited to manipulate class or concept representations. To address this, we propose an evaluation framework to assess model reliability by analyzing responses to global and local perturbations in the embedding space, enabling the identification of inputs that trigger unreliable or biased behavior. Beyond social implications, fairness and diversity are fundamental to defining robust and trustworthy model behavior. Our approach offers deeper insights into these essential aspects by evaluating: (i) generative diversity, measuring the breadth of visual representations for learned concepts, and (ii) generative fairness, which examines the impact that removing concepts from input prompts has on control, under a low guidance setup. Beyond these evaluations, our method lays the groundwork for detecting unreliable, bias-injected models and tracing the provenance of embedded biases. Our code is publicly available at https://github.com/JJ-Vice/T2I_Fairness_Diversity_Reliability. Keywords: Fairness, Reliability, AI Ethics, Bias, Text-to-Image Models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。