构建首个图文生成安全评估基准,覆盖毒性、公平性与隐私三方面风险。
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
- 设计12个任务44类别的安全评估体系,涵盖毒性、偏见与隐私风险。
- 收集7万条提示词,构建6.8万张人工标注图像数据集,发现模型普遍存在种族不公平问题。
- 开发新型评估器,可检测大模型未识别的高危内容,适合研究者与开发者使用。
文本到图像(T2I)模型快速发展,可从文本提示生成高质量跨领域图像。但这类模型存在显著安全风险,包括生成有害、偏颇或涉及隐私的内容。当前对T2I安全性的评估仍处于初级阶段,多数研究仅覆盖特定维度,诸多关键风险尚未被探索。为此,我们提出T2ISafety,一个涵盖毒性、公平性与隐私三大核心领域的安全评估基准。基于此,构建了包含12个任务和44个类别的详细分类体系,并精心收集7万条对应提示词。基于该分类体系与提示集,建立包含6.8万张人工标注图像的大规模T2I数据集,并训练出能识别先前研究未能发现的关键风险的评估器,甚至可检测连超大规模专有模型如GPTs也无法正确识别的风险。我们在T2ISafety上评估了12个主流扩散模型,揭示了持久存在的种族公平性问题、生成有毒内容的倾向,以及即使采用概念擦除等防御方法,各模型在隐私保护上仍存在显著差异。数据集与评估器已开源:https://github.com/adwardlee/t2i_safety。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have rapidly advanced, enabling the generation of high-quality images from text prompts across various domains. However, these models present notable safety concerns, including the risk of generating harmful, biased, or private content. Current research on assessing T2I safety remains in its early stages. While some efforts have been made to evaluate models on specific safety dimensions, many critical risks remain unexplored. To address this gap, we introduce T2ISafety, a safety benchmark that evaluates T2I models across three key domains: toxicity, fairness, and bias. We build a detailed hierarchy of 12 tasks and 44 categories based on these three domains, and meticulously collect 70K corresponding prompts. Based on this taxonomy and prompt set, we build a large-scale T2I dataset with 68K manually annotated images and train an evaluator capable of detecting critical risks that previous work has failed to identify, including risks that even ultra-large proprietary models like GPTs cannot correctly detect. We evaluate 12 prominent diffusion models on T2ISafety and reveal several concerns including persistent issues with racial fairness, a tendency to generate toxic content, and significant variation in privacy protection across the models, even with defense methods like concept erasing. Data and evaluator are released under https://github.com/adwardlee/t2i_safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。