arXiv:2510.22300cs.CRcs.AI2025-10AAAI被引 14

构建首个细粒度文本生成图像安全评测基准,覆盖6大类14子类风险提示。

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

  • 建立六类14子类风险分类体系,实现提示词的层级化标注。
  • 收集6432条有效风险提示,每条标注类别与具体风险原因。
  • 支持攻击、防御、检测全链条评估,适配安全研究者使用。

利用色情、暴力等风险文本提示测试文本到图像(T2I)模型的安全性至关重要。然而现有风险提示数据集在三个关键方面受限:1)风险类别有限,2)标注粗略,3)有效性不足。为此,我们提出T2I-RiskyPrompt,一个面向T2I模型安全评估的综合性基准。首先,构建包含6个主类和14个子类的层级风险分类体系;在此基础上,设计采集与标注流程,最终获得6,432条有效风险提示,每条均附带层级标签与详细风险原因。为便于评估,提出基于原因驱动的危险图像检测方法,显式对齐多模态大模型与安全标注。基于该基准,我们对8个T2I模型、9种防御方法、5个安全过滤器和5种攻击策略进行了全面评估,揭示了九项关于T2I模型安全性的关键洞见。最后讨论其在多个研究领域的应用潜力。数据集与代码已开源:https://github.com/datar001/T2I-RiskyPrompt。

原文摘要 · Abstract (English)

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To address these limitations, we introduce T2I-RiskyPrompt, a comprehensive benchmark designed for evaluating safety-related tasks in T2I models. Specifically, we first develop a hierarchical risk taxonomy, which consists of 6 primary categories and 14 fine-grained subcategories. Building upon this taxonomy, we construct a pipeline to collect and annotate risky prompts. Finally, we obtain 6,432 effective risky prompts, where each prompt is annotated with both hierarchical category labels and detailed risk reasons. Moreover, to facilitate the evaluation, we propose a reason-driven risky image detection method that explicitly aligns the MLLM with safety annotations. Based on T2I-RiskyPrompt, we conduct a comprehensive evaluation of eight T2I models, nine defense methods, five safety filters, and five attack strategies, offering nine key insights into the strengths and limitations of T2I model safety. Finally, we discuss potential applications of T2I-RiskyPrompt across various research fields. The dataset and code are provided in https://github.com/datar001/T2I-RiskyPrompt.

安全评测文本生成风险提示图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。