arXiv:2503.00020cs.CLcs.AI2025-03中稿 · publication in IEE…综述被引 5

梳理文本生成图像安全研究常用数据集,助你选对数据、避坑伦理风险。

A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety

  • 系统分析10+个T2I安全研究常用数据集的构建方式与内容构成
  • 揭示各数据集在有害内容覆盖与分布上的差异,反映真实安全风险缺口
  • 为研究人员提供数据选型指南,助力模型安全评估与未来研究设计

针对文本生成图像(T2I)生成式AI安全的新研究,常依赖公开数据集进行训练与评估,因此数据集的质量与构成至关重要。本文全面回顾了当前T2I研究中使用的关键数据集,详细描述其收集方法、数据构成、提示词的语义与句法多样性,以及有害内容的质量、覆盖范围与分布情况。通过揭示各数据集的优势与局限,本研究帮助研究者根据具体应用场景选择最相关数据集,批判性评估其工作在下游带来的影响,尤其在模型安全与伦理层面,并识别出未来研究可填补的数据集覆盖与质量空白。

原文摘要 · Abstract (English)

Novel research aimed at text-to-image (T2I) generative AI safety often relies on publicly available datasets for training and evaluation, making the quality and composition of these datasets crucial. This paper presents a comprehensive review of the key datasets used in the T2I research, detailing their collection methods, compositions, semantic and syntactic diversity of prompts and the quality, coverage, and distribution of harm types in the datasets. By highlighting the strengths and limitations of the datasets, this study enables researchers to find the most relevant datasets for a use case, critically assess the downstream impacts of their work given the dataset distribution, particularly regarding model safety and ethical considerations, and also identify the gaps in dataset coverage and quality that future research may address.

T2I安全数据集评测生成式AI伦理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。