arXiv:2412.00064cs.CVcs.AI2024-12被引 9

提出文本安全检测新方法,有效拦截扩散模型生成违规内容

DiffGuard: Text-Based Safety Checker for Diffusion Models

  • 基于文本特征构建安全过滤器,替代传统图像检测
  • 在测试中比现有最佳方案准确率提升14%以上
  • 适合需要内容安全的AI图像生成系统开发者使用

扩散模型的快速发展使得文本生成图像成为可能,以DALL-E和Midjourney为代表的闭源模型引领技术前沿。而Stable Diffusion等开源模型在Hugging Face平台也具备类似能力,并内置伦理过滤机制以防止生成不当内容。本文首次揭示了现有过滤机制的局限性,并提出一种新型文本级安全检查方法DiffGuard。该方法通过分析输入提示词中的潜在风险特征,实现更精准的内容防护。实验表明,DiffGuard在多项评估中表现优于当前最优过滤方案,性能提升超过14%,显著增强对生成内容的可控性与安全性,尤其适用于防范信息战等恶意用途。

原文摘要 · Abstract (English)

Recent advances in Diffusion Models have enabled the generation of images from text, with powerful closed-source models like DALL-E and Midjourney leading the way. However, open-source alternatives, such as StabilityAI's Stable Diffusion, offer comparable capabilities. These open-source models, hosted on Hugging Face, come equipped with ethical filter protections designed to prevent the generation of explicit images. This paper reveals first their limitations and then presents a novel text-based safety filter that outperforms existing solutions. Our research is driven by the critical need to address the misuse of AI-generated content, especially in the context of information warfare. DiffGuard enhances filtering efficacy, achieving a performance that surpasses the best existing filters by over 14%.

扩散模型内容安全文本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。