首个支持多模态的智能安全防护系统,能主动推理风险。
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- 用主动推理机制统一处理文本、图像、视频和音频的安全问题
- 在15个基准上表现优异,覆盖广泛多模态安全场景
- 适合需要跨模态安全防护的AI系统开发者
能处理文本、图像、视频和音频的多模态大语言模型(OLLM)带来了人机交互中的新安全挑战。以往的安全防护研究主要针对单模态场景,通常将防护视为二分类任务,难以在多种模态和任务间保持鲁棒性。为此,我们提出OmniGuard,首个具备主动推理能力的多模态安全防护框架,可统一处理所有模态的安全问题。为训练OmniGuard,我们构建了一个包含超过21万条样本的综合性多模态安全数据集,涵盖单模态与跨模态输入,每条样本配有结构化安全标签及专家模型通过定向蒸馏生成的安全评注。在15个基准上的实验表明,OmniGuard在多种多模态安全场景中均表现出强有效性与泛化能力。重要的是,OmniGuard提供统一框架,可执行策略并降低多模态风险,为构建更鲁棒、更强的多模态防护系统铺平道路。
原文摘要 · Abstract (English)
Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardrail research largely targets unimodal settings and typically frames safeguarding as binary classification, which limits robustness across diverse modalities and tasks. To address this gap, we propose OmniGuard, the first family of omni-modal guardrails that performs safeguarding across all modalities with deliberate reasoning ability. To support the training of OMNIGUARD, we curate a large, comprehensive omni-modal safety dataset comprising over 210K diverse samples, with inputs that cover all modalities through both unimodal and cross-modal samples. Each sample is annotated with structured safety labels and carefully curated safety critiques from expert models through targeted distillation. Extensive experiments on 15 benchmarks show that OmniGuard achieves strong effectiveness and generalization across a wide range of multimodal safety scenarios. Importantly, OmniGuard provides a unified framework that enforces policies and mitigates risks in omni-modalities, paving the way toward building more robust and capable omnimodal safeguarding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。