arXiv:2504.01081cs.CVcs.CL2025-04被引 30

ShieldGemma 2可精准识别合成与自然图像中的色情、暴力等风险内容。

ShieldGemma 2: Robust and Tractable Image Content Moderation

  • 基于Gemma 3构建40亿参数模型,支持多类安全风险预测
  • 在内外部基准上优于LlavaGuard、GPT-4o mini等主流模型
  • 开源工具助力多模态安全,适合内容审核与AI治理场景

我们提出ShieldGemma 2,一个基于Gemma 3的40亿参数图像内容安全检测模型。该模型能对合成图像(如任意生成模型输出)和自然图像(如任意视觉-语言模型输入)中的色情、暴力血腥及危险内容提供鲁棒的安全风险预测。在内部与外部基准上的评估表明,其性能优于LlavaGuard、GPT-4o mini及基础版Gemma 3模型。此外,我们还设计了一种新型对抗性数据生成流水线,实现可控、多样且鲁棒的图像生成。ShieldGemma 2作为开源图像审核工具,旨在推动多模态安全与负责任的AI发展。

原文摘要 · Abstract (English)

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore, and Dangerous Content for synthetic images (e.g. output of any image generation model) and natural images (e.g. any image input to a Vision-Language Model). We evaluated on both internal and external benchmarks to demonstrate state-of-the-art performance compared to LlavaGuard \citep{helff2024llavaguard}, GPT-4o mini \citep{hurst2024gpt}, and the base Gemma 3 model \citep{gemma_2025} based on our policies. Additionally, we present a novel adversarial data generation pipeline which enables a controlled, diverse, and robust image generation. ShieldGemma 2 provides an open image moderation tool to advance multimodal safety and responsible AI development.

内容审核多模态安全图像生成AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。