arXiv:2608.27548cs.AI2026-08

40亿参数多模态安全模型,支持12语言实时内容审核

Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator

  • 4B小模型统一处理文本、图像与回复的多模态安全判断
  • 在12种语言中实现高覆盖率且延迟敏感场景下表现优秀
  • 支持自定义策略推理,适合需可解释性审计的部署场景

部署中的AI应用安全审核正从纯文本扩展至图像、文档、截图及生成回复等多模态内容,且需适配跨领域差异化政策。现有防护机制通常覆盖有限,难以兼顾全面性、自定义策略与低算力开销。本文提出Nemotron 3.5 Content Safety Moderator(简称Nemotron 3.5 CS),一个40亿参数的紧凑型多模态多语言视觉-语言安全审核模型,可联合分类用户提示、图像与助手回复,支持12种语言。该模型能返回低延迟的安全标签,并在需要时生成简明推理轨迹,以应用自定义策略并识别违规类别。我们还发布了一个涵盖真实图像标注、良性图文/文档任务、合成罕见风险与越狱案例、以及自定义策略示例的多模态多语言安全数据集。在多模态安全、文本审核、多语言鲁棒性、自定义策略遵循、良性误报率和延迟等多项评估中,Nemotron 3.5 CS展现出实用的权衡:在保持广泛覆盖的同时,引入图像与策略条件化审核,性能仍优于或媲美专用安全模型。结果表明,紧凑型视觉-语言审核器可作为可部署的一线安全组件,推理功能则可按需用于审计与政策审查。

原文摘要 · Abstract (English)

Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this setting, making it difficult to combine broad coverage, custom policy control, and low compute cost. We present Nemotron 3.5 Content Safety Moderator, also referred to as Nemotron 3.5 CS in this paper for brevity, a compact 4B vision-language safety moderator that jointly classifies user prompts, images, and assistant responses across 12 languages. Nemotron 3.5 CS returns safety labels for latency-sensitive moderation and can additionally produce concise reasoning traces that apply supplied custom policies and identify violated categories when reasoning is requested. We also release a multimodal and multilingual safety dataset for guard training, spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples. Across evaluations spanning multimodal safety, text moderation, multilingual robustness, custom-policy following, benign false positives, and latency, Nemotron 3.5 CS demonstrates a practical coverage tradeoff: it adds image-conditioned and policy-conditioned moderation while remaining broadly competitive with specialized guard models. These results suggest that compact vision-language moderators can serve as deployable front-line safety components, with reasoning used selectively for audit and policy review.

内容安全多模态轻量模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。